Evaluating and Controlling Natural Language Inference over Clinical Trial Texts
Abstract
This thesis formalises Clinical Trial Natural Language Inference as deciding whether clinically meaningful claims are supported, contradicted, or indeterminate across heterogeneous trial materials. It develops a reasoning-first programme connecting task specification, robustness-centred evaluation, and controlled system design. The contributions span the NLI4CT benchmark and shared tasks, Faithfulness and Consistency metrics, ontology-grounded retrieval, prompt and LoRA studies, GKMRV diagnostic probes, and the CARENLI agentic framework. Collectively, the studies show how aggregate scores can conceal shortcutting and brittleness, while structured inference procedures make errors easier to locate and can improve reasoning fidelity.