Armin Toroghi

prof_pic.jpg

I am a postdoctoral researcher at the D3M lab, department of Computer Science, University of Toronto, working under the mentorship of Professor Scott Sanner. I am also a Faculty Affiliate researcher at the Vector Institute of AI and the Institute of Information & Communications Technology Planning & Evaluation. My research lies at the intersection of reasoning, LLMs, Knowledge Graphs, and Recommender Systems.

I completed my Ph.D. degree at the University of Toronto in 2025. During my Ph.D. studies, I was a student researcher collaborating with the LG Electronics AI research, and worked as an ML research intern at Borealis AI and Vector Institute of AI. Prior to joining the UofT, I completed my M.Sc. and B.Sc. at the University of Tehran, where my research was focused on developing deep learning-based methods for studying mechanical systems.

Current Research Focus: My research focuses on natural language processing and machine learning, with a particular emphasis on reasoning of large language models (LLMs). I’m interested in enhancing the reasoning capabilities of LLMs, and building agents that are more reliable, efficient, and verifiable, aimed at real-world impact in automated decision support systems and knowledge-intensive tasks.

  • Developing Efficient and Verifiable LLM Reasoning Frameworks: We develop efficient and verifiable reasoning frameworks for large language models that go beyond scale-driven performance gains. Our research focuses on improving mathematical and logical reasoning through structured training signals, uncertainty-aware inference, and principled data selection. A central motivation of our work is verifiability: as LLMs become more capable, their errors increasingly stem from confident hallucinations rather than lack of fluency. Verifiable reasoning mechanisms are essential for enabling the safe and effective use of such potent models by exposing uncertainty, enabling checks on intermediate reasoning, and supporting downstream validation.

  • Understanding and Mitigating Failure Modes of LLMs: We systematically study where and why large language models fail, with particular focus on unreliable reasoning and confident hallucinations that undermine trust in increasingly capable systems. Our work aims to move beyond surface-level accuracy by analyzing how and why models arrive at their answers, and by developing methods that encourage alignment between correct outputs and correct reasoning processes. Projects such as CoLoTa investigate failure modes in reasoning under distribution shift and limited supervision, while Right for the Right Reasons proposes solutions to the issue of unfaithful answer generation and commonsense reasoning errors.

  • Developing Methodologies for Complex Reasoning over Knowledge Graphs: We develop efficient methods for complex reasoning over knowledge graphs, which provide structured, explicit semantics that complement the generative strengths of large language models. In the LLM era, knowledge graphs play a crucial role in enabling grounded, multi-hop, and verifiable reasoning that cannot be reliably achieved through unstructured generation alone. Our projects aim to advance this direction by designing scalable approaches for expressive reasoning over large knowledge graphs.

news

Apr 02, 2026 Our paper “CommonWhy: A Dataset for Evaluating Entity-Based Causal Commonsense Reasoning in Large Language Models” has been accepted to SIGIR 2026.
Mar 11, 2026 I’m giving a talk on “Limitations of RLVR for LLM Reasoning Enhancement” at the Vector Institute of AI. Join via Zoom
Jan 26, 2026 Our paper “Natural Language PDDL (NL-PDDL) for Open-world Goal-oriented Commonsense Regression Planning in Embodied AI” has been accepted to ICLR 2026.
Dec 01, 2025 I joined the D3M lab as a postdoctoral researcher.
Nov 04, 2025 I defended my Ph.D. thesis.

selected publications

  1. LLM-based Typed Hyperresolution for Commonsense Reasoning with Knowledge Bases
    Armin Toroghi, Ali Pesaranghader, Tanmana Sadhu, and 1 more author
    In The Thirteenth International Conference on Learning Representations, 2025
  2. Verifiable, debuggable, and repairable commonsense logical reasoning via llm-based theory resolution
    Armin Toroghi, Willis Guo, Ali Pesaranghader, and 1 more author
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024
  3. Right for right reasons: Large language models for verifiable commonsense knowledge graph question answering
    Armin Toroghi, Willis Guo, Mohammad Mahdi Abdollah Pour, and 1 more author
    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024
  4. CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail Knowledge
    Armin Toroghi, Willis Guo, and Scott Sanner
    In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025
  5. Bayesian inference with complex knowledge graph evidence
    Armin Toroghi and Scott Sanner
    In Proceedings of the AAAI Conference on Artificial Intelligence, 2024
  6. Bayesian knowledge-driven critiquing with indirect evidence
    Armin Toroghi, Griffin Floto, Zhenwei Tang, and 1 more author
    In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2023
  7. Natural Language PDDL (NL-PDDL) for Open-world Goal-oriented Commonsense Regression Planning in Embodied AI
    Xiaotian Liu*, Armin Toroghi*, Jiazhou Liang, and 7 more authors
    In The Fourteenth International Conference on Learning Representations, 2026