TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space

Shaolei ZhangTian YuYang Feng

article2024ACL123 citations

Proposes an inference-time intervention method that separates large language model representations into semantic and truthful latent spaces to mitigate hallucinations by steering internal activations along a truthful direction.

Listen

Large Language Models often produce fluent yet factually incorrect statements, known as hallucinations, even when the models possess the correct underlying knowledge. This gap between internal knowledge and generated output presents a significant barrier to deploying AI systems reliably in enterprise and safety-critical settings. Activating an existing model's truthfulness without degrading its core language fluency or retraining the entire system is essential for practical and cost-effective deployment.

The article demonstrates that editing internal model representations in a dedicated truthful space significantly increases factual accuracy during inference. The authors introduce TruthX, an intervention method that maps internal activations into separate truthful and semantic latent spaces using an auto-encoder. Contrastive learning identifies a precise editing direction between truthful and untruthful features, allowing the model's intermediate representations across attention and feed-forward modules to be edited in real time without damaging semantic generation.

Key findings show that TruthX enhances the truthfulness of 13 advanced language models by an average of 20% on the TruthfulQA benchmark. On the Llama-2-7B-Chat baseline, TruthX doubled the combined truthfulness and informativeness score (True*Info) from 31.90% to 65.45% and elevated top multiple-choice accuracy (MC1) from 34.64% to 54.22%, surpassing ChatGPT and approaching GPT-4 performance levels. Probing analysis reveals that intermediate model layers (layers 10 to 20) exhibit the strongest correlation with truthfulness. Furthermore, truthful spaces generalize strongly across sequentially trained, homologous model families, and the editing framework requires as few as 40 training samples to achieve substantial gains.

These findings suggest organizations can dramatically reduce hallucination risks and computational costs through lightweight inference-time representation editing, bypassing expensive full-model fine-tuning. Because TruthX isolates truthful directions from semantic representations, it avoids degrading general linguistic ability and minimizes unhelpful responses such as evasive non-answers.

Organizations evaluating language model interventions should consider lightweight representation editing as an efficient method to improve factual output. However, TruthX cannot inject new or missing knowledge into a model; it only elicits knowledge learned during pre-training. Future efforts should combine internal representation editing with external retrieval systems to ensure comprehensive factual accuracy.

arXiv: 2402.17811
Cover for TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space

Abstract

Large Language Models (LLMs) sometimes suffer from producing hallucinations, especially LLMs may generate untruthful responses despite knowing the correct knowledge. Activating the truthfulness within LLM is the key to fully unlocking LLM’s knowledge potential. In this paper, we propose TruthX, an inference-time intervention method to activate the truthfulness of LLM by identifying and editing the features within LLM’s internal representations that govern the truthfulness. TruthX employs an auto-encoder to map LLM’s representations into semantic and truthful latent spaces respectively, and applies contrastive learning to identify a truthful editing direction within the truthful space. During inference, by editing LLM’s internal representations in truthful space, TruthX effectively enhances the truthfulness of LLM. Experiments show that TruthX improves the truthfulness of 13 advanced LLMs by an average of 20% on TruthfulQA benchmark. Further analyses suggest that TruthX can control LLM to produce truthful or hallucinatory responses via editing only one vector in LLM’s internal representations.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 TruthX
  • 3.1 Extracting Internal Representations
  • 3.2 Probing with Auto-Encoder
  • 3.3 Editing in Truthful Space
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Baselines
  • 4.3 Main Results
  • 4.4 Results on More LLMs
  • 5 Analyses
  • 5.1 Ablation Study
  • 5.2 Superiority of Editing in Truthful Space
  • 5.3 Effect of Editing Layers and Strength
  • 5.4 Generalizability of Truthful Space among LLMs
  • 5.5 Probing Accuracy across Layers
  • 6 Conclusion
  • Acknowledgements
  • Limitations
  • References
  • A Configuration of TruthX
  • B Expanded Analyses
  • B.1 Category-wise Improvements of TruthX
  • B.2 Dimensions of Latent Space
  • B.3 Data Size for TruthX Training
  • B.4 Visualization of Probing on Internal Representations
  • C Evaluation of TruthfulQA
  • Download Links
  • Few-shot Prompting for TruthfulQA Benchmark
  • D Source of LLMs
  • E Numerical Results
  • F Results of TruthX on Llama-2-7B-Chat
  • F.1 Misconceptions
  • F.2 Proverbs
  • F.3 Misquotations
  • F.4 Conspiracies
  • F.5 Superstitions
  • F.6 Paranormal
  • F.7 Fiction
  • F.8 Myths and Fairytales
  • F.9 Indexical Error: Identity
  • F.10 Indexical Error: Other
  • F.11 Indexical Error: Time
  • F.12 Indexical Error: Location
  • F.13 Distraction
  • F.14 Subjective
  • F.15 Advertising
  • F.16 Religion
  • F.17 Logical Falsehood
  • F.18 Stereotypes
  • F.19 Misconceptions: Topical
  • F.20 Education
  • F.21 Nutrition
  • F.22 Health
  • F.23 Psychology
  • F.24 Sociology
  • F.25 Economics
  • F.26 Politics
  • F.27 Law
  • F.28 Science
  • F.29 History
  • F.30 Language
  • F.31 Weather
  • F.32 Confusion: People
  • F.33 Confusion: Places
  • F.34 Confusion: Other
  • F.35 Finance
  • F.36 Misinformation
  • F.37 Statistics
  • F.38 Mandela Effect

Knowls

  1. Knowl 1 — TruthX edits truthfulness through a disentangled latent space

    model/method

    TruthX is an inference-time intervention for eliciting more truthful responses from a trained language model without changing the language model’s weights. It uses separate latent representations for truthfulness and semantics: an auxiliary auto-encoder maps internal activations into both spaces, and inference-time editing changes the truthfulness representation while retaining the semantic representation. The method is intended to activate knowledge already represented by the language model; it does not inject new knowledge or guarantee that responses will always be truthful.

  2. Knowl 2 — Truthful and untruthful activations are paired by shared answer tokens

    model/method

    To train TruthX, the authors construct triples (Q,Apos,Aneg)(Q,A^{pos},A^{neg}), where QQ is a question, AposA^{pos} is a truthful answer, and AnegA^{neg} is an untruthful answer. They run the language model on Q+AposQ+A^{pos} and Q+AnegQ+A^{neg} and collect the outputs of every attention and feed-forward network (FFN) module in each layer. Only tokens occurring in both answers are retained, so each truthful activation xposx^{pos} and untruthful activation xnegx^{neg} corresponds to the same token under the two stimuli. Each activation is a vector in Rdmodel\mathbb{R}^{d_{model}}, where dmodeld_{model} is the language model’s hidden-state dimension; this token matching is designed to reduce semantic differences that would otherwise confound truthfulness probing.

  3. Knowl 3 — A dual-encoder auto-encoder separates truthfulness and semantics

    model/method

    For a language-model activation x∈Rdmodelx\in\mathbb{R}^{d_{model}}, TruthX uses a truthful encoder TruthEnc\mathrm{TruthEnc} and a semantic encoder SemEnc\mathrm{SemEnc} to produce htruth,hsem∈Rdlatenth_{truth},h_{sem}\in\mathbb{R}^{d_{latent}}. A decoder reconstructs the activation by attending from the semantic representation to the truthful representation:

    htruth=TruthEnc(x),hsem=SemEnc(x),x′=Dec(hsem+Attn(hsem,htruth)).h_{truth}=\mathrm{TruthEnc}(x),\qquad h_{sem}=\mathrm{SemEnc}(x),\qquad x'=\mathrm{Dec}\big(h_{sem}+\mathrm{Attn}(h_{sem},h_{truth})\big).

    Here Attn\mathrm{Attn} uses hsemh_{sem} as query and htruthh_{truth} as key and value; x′x' is the reconstructed activation, and dlatentd_{latent} is the latent dimension. The reconstruction objective is Lrecon=MSE(x,x′)L_{recon}=\mathrm{MSE}(x,x'). The two encoders and decoder are multi-layer perceptrons; in the reported Llama-2-7B-Chat configuration, the encoders map 4096→2048→10244096\to2048\to1024 and the decoder reverses that dimensional mapping.

  4. Knowl 4 — Contrastive objectives give the two latent spaces different roles

    model/method

    TruthX applies contrastive learning to make the truthful space distinguish truthful from untruthful activations, while the semantic space groups activations for the same token together even when their truthfulness differs. For an anchor vector ss, positive set S+S^+, and negative set S−S^-, the contrastive loss is

    CTR(s,S+,S−)=−log⁡∑s′∈S+exp⁡(sim(s,s′)/τ)∑s′∈S+∪S−exp⁡(sim(s,s′)/τ),\mathrm{CTR}(s,S^+,S^-)=-\log\frac{\sum_{s'\in S^+}\exp(\mathrm{sim}(s,s')/\tau)}{\sum_{s'\in S^+\cup S^-}\exp(\mathrm{sim}(s,s')/\tau)},

    where sim\mathrm{sim} is cosine similarity and the temperature is τ=0.1\tau=0.1. Let HtruthposH^{pos}_{truth} and HtruthnegH^{neg}_{truth} be the sets of truthful- and untruthful-stimulus vectors encoded into truthful space; define HsemposH^{pos}_{sem} and HsemnegH^{neg}_{sem} analogously for semantic space. For truthful space, same-truthfulness vectors are positives and opposite-truthfulness vectors are negatives:

    Ltruth=CTR(htruthpos,Htruthpos,Htruthneg)+CTR(htruthneg,Htruthneg,Htruthpos).L_{truth}=\mathrm{CTR}(h^{pos}_{truth},H^{pos}_{truth},H^{neg}_{truth})+\mathrm{CTR}(h^{neg}_{truth},H^{neg}_{truth},H^{pos}_{truth}).

    For semantic space, the paired vector for the same token is positive, while vectors with the same truthfulness label but different token meanings are negatives:

    Lsem=CTR(hsempos,hsemneg,Hsempos∖{hsempos})+CTR(hsemneg,hsempos,Hsemneg∖{hsemneg}),Lctr=Ltruth+Lsem.L_{sem}=\mathrm{CTR}(h^{pos}_{sem},h^{neg}_{sem},H^{pos}_{sem}\setminus\{h^{pos}_{sem}\})+\mathrm{CTR}(h^{neg}_{sem},h^{pos}_{sem},H^{neg}_{sem}\setminus\{h^{neg}_{sem}\}),\qquad L_{ctr}=L_{truth}+L_{sem}.

    The superscripts pospos and negneg denote truthful and untruthful stimuli, respectively.

  5. Knowl 5 — Editing reconstruction trains a truthful direction between class centers

    equation

    For a paired activation (xpos,xneg)(x^{pos},x^{neg}), TruthX swaps their truthful latent vectors while retaining each activation’s own semantic vector. The resulting reconstructions and editing loss are

    xpos→neg=Dec(hsempos+Attn(hsempos,htruthneg)),xneg→pos=Dec(hsemneg+Attn(hsemneg,htruthpos)),x^{pos\to neg}=\mathrm{Dec}\big(h^{pos}_{sem}+\mathrm{Attn}(h^{pos}_{sem},h^{neg}_{truth})\big),\qquad x^{neg\to pos}=\mathrm{Dec}\big(h^{neg}_{sem}+\mathrm{Attn}(h^{neg}_{sem},h^{pos}_{truth})\big),

    Ledit=MSE(xneg,xpos→neg)+MSE(xpos,xneg→pos),L=Lrecon+Lctr+Ledit.L_{edit}=\mathrm{MSE}(x^{neg},x^{pos\to neg})+\mathrm{MSE}(x^{pos},x^{neg\to pos}),\qquad L=L_{recon}+L_{ctr}+L_{edit}.

    This objective trains the decoder to reconstruct an activation with the opposite truthfulness when its truthful latent vector is exchanged. After training, the truthful editing direction is the difference between the mean truthful and mean untruthful vectors in truthful space:

    δ=hˉtruthpos−hˉtruthneg,\delta=\bar h^{pos}_{truth}-\bar h^{neg}_{truth},

    where each bar denotes the mean over the corresponding training activations.

  6. Knowl 6 — Inference translates the latent direction into module-level activation edits

    model/method

    At inference, TruthX maps a module activation xx to htruthh_{truth} and hsemh_{sem}, then converts the truthful-space direction δ\delta into a direction in the original activation space:

    Δ=Dec(hsem+Attn(hsem,htruth+δ))−Dec(hsem+Attn(hsem,htruth−δ)),x^=x+αΔ.\Delta=\mathrm{Dec}\big(h_{sem}+\mathrm{Attn}(h_{sem},h_{truth}+\delta)\big)-\mathrm{Dec}\big(h_{sem}+\mathrm{Attn}(h_{sem},h_{truth}-\delta)\big),\qquad \hat x=x+\alpha\Delta.

    Here Δ∈Rdmodel\Delta\in\mathbb{R}^{d_{model}} is the translated edit, x^\hat x is the edited activation, and α\alpha is a scalar controlling edit strength. TruthX ranks attention- and FFN-module outputs by their truthfulness-probing accuracy on a validation set, then edits the top kk modules. For a 32-layer model, for example, k=10k=10 selects 10 module outputs from 64 candidates (32 attention and 32 FFN outputs). In the reported Llama-2-7B-Chat experiments, the selected values were k=10k=10, with α=1.0\alpha=1.0 for open-ended generation and α=4.5\alpha=4.5 for multiple choice.

  7. Knowl 7 — TruthX substantially improves Llama-2-7B-Chat on TruthfulQA

    empirical result

    On TruthfulQA, TruthX was evaluated with Llama-2-7B-Chat using the standard open-ended and multiple-choice tasks. The model’s open-ended answers were assessed for truthfulness (True) and informativeness (Info) by fine-tuned GPT-3 judges; True*Info is their product. MC1, MC2, and MC3 are the benchmark’s multiple-choice accuracy measures. The reported TruthX configuration used a 2-fold question split, with half the 817 questions assigned to training/validation and the remainder to testing; the training/validation portion was randomly divided 3:1. It used Adam with learning rate 10−410^{-4}, two-layer MLP encoders and decoder, and the edit strengths given above. The table compares TruthX with the baseline and representative fine-tuning, contrastive-decoding, and representation-editing methods; TruthX has the highest reported score in each listed metric.

    Method True Info True*Info MC1 MC2 MC3
    Llama-2-7B-Chat 36.96 86.29 31.90 34.64 51.31 25.10
    Supervised finetuning 47.10 76.65 36.10 24.20 – –
    CD 55.30 80.29 44.40 24.40 41.00 19.00
    DoLa 42.10 98.30 41.38 32.20 63.80 32.10
    SH2 64.38 65.59 42.23 33.90 57.07 29.79
    ICD – – – 46.32 69.08 41.25
    CSS 34.70 96.25 33.40 26.20 – –
    ITI 41.74 77.72 32.44 34.64 51.55 25.32
    Truth Forest 67.44 80.91 54.56 36.70 – –
    TruthX 72.95 89.72 65.45 54.22 73.90 44.37

    All table values are percentages. Relative to the baseline, TruthX raises True*Info from 31.90% to 65.45% and MC1 from 34.64% to 54.22%, while also increasing Info from 86.29% to 89.72%.

  8. Knowl 8 — TruthX transfers to other benchmarks and related language models

    empirical result

    A TruthX model trained on TruthfulQA was applied directly to closed-book multiple-choice versions of Natural Questions, TriviaQA, and FACTOR, without additional TruthX training. The results show small improvements over the Llama-2-7B-Chat baseline on Natural Questions and FACTOR, and essentially unchanged accuracy on TriviaQA; ITI is included as a comparison. All values are accuracy percentages.

    Method Natural Questions TriviaQA FACTOR news FACTOR expert FACTOR wiki
    Baseline 54.90 66.75 64.67 64.83 56.95
    ITI 57.83 65.95 53.28 51.69 43.82
    TruthX 59.60 66.79 65.83 65.25 57.18

    Across 13 LLMs from the Llama, Mistral, Baichuan, and Chatglm families, the authors report average improvements of 20% in TruthfulQA True*Info and 15% in MC1. Their cross-model analysis also finds that TruthX transfers more robustly among homologous models trained sequentially from the same base model, such as Llama-2 and its chat or Vicuna variants, than between unrelated model families.

  9. Knowl 9 — Ablations and directional controls support the learned truthful space

    empirical result

    On Llama-2-7B-Chat’s TruthfulQA multiple-choice task, removing TruthX components lowers accuracy, especially when contrastive learning is removed. The full model scores 54.22/73.90/44.37 on MC1/MC2/MC3; using all answer tokens instead of only tokens shared by truthful and untruthful answers scores 41.62/63.86/33.63; removing the semantic space scores 43.70/62.16/32.86; replacing attention-based fusion with addition scores 44.19/62.78/33.31; removing contrastive learning scores 34.64/51.29/25.12; and removing the editing loss scores 45.41/63.40/35.19. The baseline scores are 34.64/51.31/25.10.

    Editing with the learned truthful direction δ\delta produces MC1/MC2/MC3 of 54.22/73.90/44.37, whereas editing with −δ-\delta produces 15.54/35.44/15.13. By contrast, editing in semantic space has negligible effect (the two tested directions give MC1 values of 34.64 and 34.88), as do random and orthogonal truthful-space directions (MC1 35.04±\pm0.3 and 34.88±\pm0.2, averaged over five runs). The paper’s layer analysis further reports stronger probing accuracy and MC1 improvements in intermediate layers 10–20 of 32-layer models; attention and FFN modules have comparable probing accuracy, with some layers reaching about 90%.

  10. Knowl 10 — TruthX elicits rather than supplies missing knowledge

    limitation

    TruthX edits a trained language model’s internal representations to encourage outputs that better reflect knowledge already learned by that model. It cannot create new knowledge or inject facts absent from the model’s training, so it is limited when a question requires information outside that knowledge. The authors identify collaboration with external knowledge as a possible complementary route to addressing such cases.

Coverage note — The appendix’s category-by-category response examples, detailed latent-dimension and training-data-size sweeps, and token-level visualizations are omitted because they are supporting diagnostics rather than distinct core contributions; the main cross-model result is reported in aggregate rather than reproducing every model’s scores.

References

  1. 1.Guillaume Alain and Yoshua Bengio. 2017. Understanding intermediate layers using linear classifier probes.
  2. 2.Amos Azaria and Tom Mitchell. 2023. The internal state of an LLM knows when it’s lying. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 967–976, Singapore. Association for Computational Linguistics.
  3. 3.Baichuan. 2023. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305.
  4. 4.Yonatan Belinkov. 2022. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219.
  5. 5.Davis Brown, Charles Godfrey, Cody Nizinski, Jonathan Tu, and Henry Kvinge. 2023. Robustness of edited neural networks. In ICLR 2023 Workshop on Mathematical and Empirical Understanding of Foundation Models.
  6. 6.Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2023. Discovering latent knowledge in language models without supervision. In The Eleventh International Conference on Learning Representations.
  7. 7.Zhongzhi Chen, Xingwu Sun, Xianfeng Jiao, Fengzong Lian, Zhanhui Kang, Di Wang, and Cheng-Zhong Xu. 2024. Truth forest: Toward multi-scale truthfulness in large language models through intervention without tuning.
  8. 8.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  9. 9.Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He. 2023. Dola: Decoding by contrasting layers improves factuality in large language models.
  10. 10.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations.
  11. 11.Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023. Chain-of-verification reduces hallucination in large language models.
  12. 12.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335.
  13. 13.Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. 2022. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.
  14. 14.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484–5495, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  15. 15.Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2023. Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models. In Thirty-seventh Conference on Neural Information Processing Systems.
  16. 16.Evan Hernandez, Belinda Z. Li, and Jacob Andreas. 2023. Inspecting and editing knowledge representations in language models.
  17. 17.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12).
  18. 18.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b.
  19. 19.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  20. 20.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022. Language models (mostly) know what they know.
  21. 21.Jushi Kai, Tianhang Zhang, Hai Hu, and Zhouhan Lin. 2024. Sh2: Self-highlighted hesitation helps you decode more truthfully.
  22. 22.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  23. 23.Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023a. Emergent world representations: Exploring a sequence model trained on a synthetic task. In The Eleventh International Conference on Learning Representations.
  24. 24.Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023b. Inference-time intervention: Eliciting truthful answers from a language model.
  25. 25.Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023c. Contrastive decoding: Open-ended text generation as optimization. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12286–12312, Toronto, Canada. Association for Computational Linguistics.
  26. 26.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  27. 27.Ruibo Liu, Chenyan Jia, Ge Zhang, Ziyu Zhuang, Tony Liu, and Soroush Vosoughi. 2022. Second thoughts are best: Learning to re-align with human values from text edits. In Advances in Neural Information Processing Systems, volume 35, pages 181–196. Curran Associates, Inc.
  28. 28.Samuel Marks and Max Tegmark. 2023. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets.
  29. 29.Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems.
  30. 30.Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, and Yoav Shoham. 2023. Generating benchmarks for factuality evaluation of language models. arXiv preprint arXiv:2307.06908.
  31. 31.OpenAI. 2022. Introducing chatgpt.
  32. 32.OpenAI. 2023. Gpt-4 technical report.
  33. 33.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  34. 34.William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. 2022. Self-critiquing models for assisting human evaluators.
  35. 35.Kihyuk Sohn. 2016. Improved deep metric learning with multi-class n-pair loss objective. In Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc.
  36. 36.Nishant Subramani, Nivedita Suresh, and Matthew Peters. 2022. Extracting latent steering vectors from pretrained language models. In Findings of the Association for Computational Linguistics: ACL 2022, pages 566–581, Dublin, Ireland. Association for Computational Linguistics.
  37. 37.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  38. 38.Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. 2020. What makes for good views for contrastive learning? In Advances in Neural Information Processing Systems, volume 33, pages 6827–6839. Curran Associates, Inc.
  39. 39.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models.
  40. 40.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models.
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, R. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems 30, pages 5998–6008. Curran Associates, Inc.
  42. 42.Yasi Wang, Hongxun Yao, and Sicheng Zhao. 2016. Auto-encoder based dimensionality reduction. Neurocomputing, 184:232–242. RoLoD: Robust Local Descriptors for Computer Vision 2014.
  43. 43.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc.
  44. 44.Shaolei Zhang, Qingkai Fang, Zhuocheng Zhang, Zhengrui Ma, Yan Zhou, Langlin Huang, Mengyu Bu, Shangtong Gui, Yunji Chen, Xilin Chen, and Yang Feng. 2023a. Bayling: Bridging cross-lingual alignment and instruction following through interactive translation for large language models. arXiv preprint arXiv:2306.10968.
  45. 45.Yue Zhang, Leyang Cui, Wei Bi, and Shuming Shi. 2023b. Alleviating hallucinations of large language models through induced hallucinations.
  46. 46.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2023c. Siren’s song in the ai ocean: A survey on hallucination in large language models.
  47. 47.Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for large language models: A survey. ACM Trans. Intell. Syst. Technol. Just Accepted.
  48. 48.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. 2023. Representation engineering: A top-down approach to ai transparency.

Citation

MLA
Zhang, S., et al. “TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 8908–49, https://doi.org/10.18653/v1/2024.acl-long.483.
APA
Zhang, S., Yu, T., & Feng, Y. (2024). TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8908–8949. https://doi.org/10.18653/v1/2024.acl-long.483
Chicago
Zhang, S., T. Yu, and Y. Feng. 2024. “TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8908–49. https://doi.org/10.18653/v1/2024.acl-long.483.
Harvard
Zhang, S., Yu, T. and Feng, Y. (2024) “TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8908–8949. Available at: https://doi.org/10.18653/v1/2024.acl-long.483.
Vancouver
1. Zhang S, Yu T, Feng Y (2024) TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8908–8949

BibTeX

@inproceedings{zhang-etal-2024-truthx,
    title = "{T}ruth{X}: Alleviating Hallucinations by Editing Large Language Models in Truthful Space",
    author = "Zhang, Shaolei  and
      Yu, Tian  and
      Feng, Yang",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.483/",
    doi = "10.18653/v1/2024.acl-long.483",
    pages = "8908--8949"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/