Reliable Retrieval-Augmented Generation
Systems and benchmarks for grounding LLM behavior in external data, knowledge graphs, and verifiable retrieval pipelines.
My research vision is to make AI systems reason reliably over data. Rooted in data management, I develop retrieval-augmented generation and agent memory that turn raw data and knowledge into trustworthy LLM and agent behavior — from question answering to scientific discovery.
I build data-centric methods and benchmarks for language models that must retrieve, remember, and reason over structured and unstructured knowledge.
Systems and benchmarks for grounding LLM behavior in external data, knowledge graphs, and verifiable retrieval pipelines.
Memory structures, freshness checks, and anchoring mechanisms for agents operating across extended conversations and tasks.
Methods for semantic typing, text-to-SQL, taxonomy reasoning, and scientific data interfaces that connect models with real data.
Tip: click a topic to filter papers · Shift-click to add more topics.
* Equal contribution + Corresponding author