CAROT

Techniques to mitigate imbalances in neurosymbolic learning

Mitigating learning imbalances (a problem typically referred to as long-tailed learning) has received considerable attention in supervised and weakly supervised learning with the proposed techniques operating at training or at testing time. However, these previous algorithms are not appropriate for neurosymbolic learning.

We propose a statistically consistent technique for estimating the marginals of the hidden labels given weak labels and two algorithms to mitigate imbalances during training and testing time. The first algorithm assigns pseudolabels to training data based on a novel linear programming formulation of neurosymbolic learning. The second algorithm uses the marginals of the hidden labels to constrain the model’s predictions on test data using robust semi-constrained optimal transport. Our empirical analysis shows that our techniques can improve the accuracy over strong baselines in neurosymbolic, Semantic Loss and Scallop, and long-tailed learning, Logic Adjustment and RECORDS by up to 14%.

Repository

Python Library for Neurosymbolic Learning Under Imbalances

Relevant publications

2025

  1. NeurIPS
    Imbalances in Neurosymbolic Learning: Characterization and Mitigating Strategies
    Efthymia Tsamoura, Kaifu Wang, and Dan Roth
    In Proceedings of the Thirty-Ninth Conference on Neural Information Processing Systems (NeurIPS), 2025