TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
Shaolei ZhangTian YuYang Feng
Proposes an inference-time intervention method that separates large language model representations into semantic and truthful latent spaces to mitigate hallucinations by steering internal activations along a truthful direction.
Large Language Models often produce fluent yet factually incorrect statements, known as hallucinations, even when the models possess the correct underlying knowledge. This gap between internal knowledge and generated output presents a significant barrier to deploying AI systems reliably in enterprise and safety-critical settings. Activating an existing model's truthfulness without degrading its core language fluency or retraining the entire system is essential for practical and cost-effective deployment.
The article demonstrates that editing internal model representations in a dedicated truthful space significantly increases factual accuracy during inference. The authors introduce TruthX, an intervention method that maps internal activations into separate truthful and semantic latent spaces using an auto-encoder. Contrastive learning identifies a precise editing direction between truthful and untruthful features, allowing the model's intermediate representations across attention and feed-forward modules to be edited in real time without damaging semantic generation.
Key findings show that TruthX enhances the truthfulness of 13 advanced language models by an average of 20% on the TruthfulQA benchmark. On the Llama-2-7B-Chat baseline, TruthX doubled the combined truthfulness and informativeness score (True*Info) from 31.90% to 65.45% and elevated top multiple-choice accuracy (MC1) from 34.64% to 54.22%, surpassing ChatGPT and approaching GPT-4 performance levels. Probing analysis reveals that intermediate model layers (layers 10 to 20) exhibit the strongest correlation with truthfulness. Furthermore, truthful spaces generalize strongly across sequentially trained, homologous model families, and the editing framework requires as few as 40 training samples to achieve substantial gains.
These findings suggest organizations can dramatically reduce hallucination risks and computational costs through lightweight inference-time representation editing, bypassing expensive full-model fine-tuning. Because TruthX isolates truthful directions from semantic representations, it avoids degrading general linguistic ability and minimizes unhelpful responses such as evasive non-answers.
Organizations evaluating language model interventions should consider lightweight representation editing as an efficient method to improve factual output. However, TruthX cannot inject new or missing knowledge into a model; it only elicits knowledge learned during pre-training. Future efforts should combine internal representation editing with external retrieval systems to ensure comprehensive factual accuracy.
- Paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, Kenneth Li et al. (2023). TruthX develops the inference-time truth steering approach introduced here, so this paper clarifies the predecessor method and its limitations.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). TruthX evaluates its gains on TruthfulQA, and this work explains the benchmark’s design and the falsehoods it measures.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). Its representation-engineering framework provides the conceptual basis for reading and steering internal model representations, the central operation in TruthX.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao 0043 et al. (2025). SPARE extends inference-time representation steering with sparse autoencoder features, applying a related editing strategy to control reliance on internal knowledge versus retrieved context.
