Tirthankar Ghosal is a staff scientist at Oak Ridge National Laboratory (ORNL), where he leads research at the intersection of AI for Science, DOE laboratory operations, high-performance computing, and quantum computing. His team develops and applies advanced AI methods to accelerate scientific discovery while improving the efficiency and reliability of HPC resources through predictive and operational analytics. As the Agentic Software project lead in the Quantum Science Center, Ghosal leads a team developing AI-enabled solutions across the quantum-HPC stack. His AI for Science research includes hypothesis generation, large language models for astrophysics and materials science, cosmological simulations, and agentic systems for scientific discovery. He also leads this research area for FORUM-AI, a multi-institutional DOE SciDAC Genesis Mission project led by Lawrence Berkeley National Laboratory, where he serves as an affiliate scientist. As principal investigator for AI for Laboratory Operations projects, Ghosal leads efforts in hazardous-event forecasting, complex engineering-diagram understanding, high-risk property classification, procurement analytics, and project-velocity modeling. He also teaches generative AI and natural language processing to advanced PhD students in the Bredesen Center at the University of Tennessee, Knoxville.
Presentation Title:
From Hypothesis Generation to Autonomous Laboratories: Trustworthy AI Co-Scientists for the Genesis Era
Presentation Abstract:
Scientific discovery is limited less by data than by our capacity to form, ground, and trust new hypotheses amid exponentially growing literature and deepening specialization. This talk presents our efforts towards treating hypothesis generation not as an isolated text task but as the first stage of an end-to-end discovery loop destined to run inside autonomous, self-driving laboratories. We first chart the landscape through comprehensive surveys of LLM-driven hypothesis generation, organizing it into prompting and fine-tuning, knowledge-enhanced retrieval, multi-agent collaboration, and reasoning-centered architectures, and naming hallucination, knowledge integration, and the novelty versus feasibility tension as its core obstacles.
We then contribute concrete methods. We reframe hypothesis generation as language generation with explicit reasoning supervision through the HypoGen dataset and a Bit-Flip-Spark schema and show that efficient state-space architectures can rival transformers on this task. Cross-domain studies across chemistry, physics, computer science, economics, and medicine map where models align with expert judgment. Moving from fluency to grounded reasoning, we introduce a redundancy-aware Compressive Knowledge Graph hypothesis, showing that compact structured subgraphs carry most of the useful signal, and Graph-PRefLexOR, graph-native models trained with reinforcement learning to make mechanism exploration and hypothesis synthesis inspectable and traceable in materials design.
Generation alone is insufficient; the laboratory of the future needs agents that act and evaluations we can trust. We connect these advances to autonomous experimentation through a modular, cross-facility architecture in which language-driven multi-agent systems orchestrate additive-manufacturing and high-performance-computing resources in near real time, and PROV-AGENT, a provenance framework making agent decisions transparent, traceable, and reproducible. Evaluation is the linchpin: Matter to Mechanism benchmarks problem-to-hypothesis reasoning in battery materials beyond surface similarity, while DISCERN, our multi-lens diagnostic, shows that a single non-triviality score conflates semantic novelty, expert excitement, and genuine epistemic openness, and tests which a model actually tracks.
Together these efforts form a coherent arc from generating hypotheses, to grounding them, to acting through accountable agents, to evaluating whether they are worth pursuing. This arc aligns with the DOE Genesis Mission, which marshals AI, supercomputing, and the national laboratory ecosystem to accelerate discovery. We close with a vision of trustworthy AI co-scientists and self-driving laboratories as a closed, auditable loop, and argue that this work marks concrete, measurable displacement toward that goal.