Updated
Updated · arxiv.org · Jul 20
SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
Updated
Updated · arxiv.org · Jul 20

SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs

1 articles · Updated · arxiv.org · Jul 20

Summary

  • Researchers have introduced SATQuest, a new verifier to evaluate and enhance the logical reasoning abilities of large language models (LLMs).
  • SATQuest generates diverse SAT-based reasoning tasks from CNF instances and objectively checks answers using PySAT, enabling fine-grained, reproducible analysis.
  • The tool exposes reasoning gaps in LLMs, especially with complex or unfamiliar formats, and demonstrates that reinforcement fine-tuning can improve logical performance.