Position: TrustLLM: Trustworthiness in Large Language Models
Yue HuangLichao SunHaoran WangSiyuan WuQihui ZhangYuan LiChujie GaoYixin HuangWenhan LyuYixuan Zhang
Establishes a standardized framework and 30-dataset benchmark across six core dimensions to rigorously evaluate the trustworthiness of 16 mainstream large language models, revealing key performance gaps between proprietary and open-source systems.
Large language models are rapidly being integrated into mission-critical domains, including software engineering, finance, healthcare, and law. However, these systems introduce serious operational, ethical, and security risks, such as generating fabricated information, leaking private data, exhibiting social biases, and falling prey to adversarial attacks. Currently, organizations lack standardized benchmarks and clear principles to evaluate these risks comprehensively before deployment.
The main objective of the article is to establish a unified evaluation framework for model trustworthiness, titled TRUSTLLM. It formulates clear principles across eight dimensions of trustworthiness and quantitatively evaluates 16 mainstream models across six practical dimensions using more than 30 benchmark datasets.
To conduct this evaluation, the authors synthesized findings from 500 studies to define eight core facets of trustworthiness: truthfulness, safety, fairness, robustness, privacy, machine ethics, transparency, and accountability. They then constructed a multi-task benchmark spanning classification and text generation across more than 30 datasets. The assessment tested 16 leading models, including prominent commercial systems such as GPT-4 and ChatGPT alongside widely accessible open-weight models such as the Llama2 family, Mistral, and Vicuna.
The investigation produced several key findings. First, a system's trustworthiness is positively correlated with its general task capability; models with superior language understanding and reasoning consistently achieve higher accuracy in moral judgment and resist adversarial inputs better. Second, commercial proprietary models generally outperform open-source counterparts in safety and reliability, though advanced open models like Llama2 demonstrate that open-weight systems can achieve comparable trustworthiness without relying on external moderators. Third, many systems suffer from exaggerated safety, improperly refusing harmless user requests; for example, Llama2-7b exhibited a 57% refusal rate on benign prompts. Fourth, performance across specific risks remains weak: all models struggle with zero-shot commonsense reasoning and internal truthfulness, identify stereotypes poorly (even GPT-4 achieved only 65% accuracy), and show varying degrees of private data leakage.
These findings indicate that deploying models based solely on standard performance benchmarks introduces unmanaged safety, regulatory compliance, and brand-reputation risks. The presence of exaggerated safety behaviors illustrates that current alignment techniques often teach superficial pattern matching rather than genuine user intent, reducing model utility. Furthermore, relying entirely on internal model knowledge creates high factual error rates, whereas augmenting systems with verified external data substantially improves truthfulness and performance.
Decision-makers and developers should adopt specific technical and organizational practices based on these results. Organizations must prioritize integrating external knowledge retrieval to mitigate hallucinations rather than relying on internal model weights. Development teams should refine alignment training using contextual intent recognition to minimize false-positive refusals. At the governance level, stakeholders across industry, academia, and the open-source community should establish shared alliances and demand transparency for alignment techniques and safety guardrails, enabling standardized independent auditing.
The findings are bounded by certain methodological constraints. The benchmark evaluates models exclusively in English, which leaves non-English safety and cultural nuances unaddressed and may disadvantage models developed in other linguistic contexts. Additionally, the evaluation relies on empirical test sets rather than mathematical guarantees, meaning it cannot certify worst-case system behavior under novel adversarial attacks. Consequently, leaders should view these findings with high confidence for standard English deployments while maintaining human oversight and domain-specific validation in high-stakes environments.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). This survey provides the foundational taxonomy of broad LLM evaluation protocols and benchmarks that TrustLLM systematically synthesizes and standardizes into its trustworthiness framework.
- Paper: Ethical and social risks of harm from Language Models, Laura Weidinger et al. (2022). It establishes the foundational categorization of social and ethical harms in language models, directly informing the core trustworthiness dimensions assessed in TrustLLM.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). This benchmark introduces standard methodologies for testing model truthfulness and imitative falsehoods, forming an essential component of TrustLLM's truthfulness evaluation.
- Paper: A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly, Yifan Yao et al. (2023). It catalogs the security, privacy, and adversarial vulnerability landscape that TrustLLM incorporates into its multi-dimensional reliability benchmarks.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). It establishes the baseline evaluation methods for measuring toxic text generation in autoregressive models, which TrustLLM adapts for safety testing.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). This paper presents the open-weight Llama 2 alignment and safety methodology that serves as a primary subject and baseline for TrustLLM's comparative safety and refusal analysis.
- Paper: On the Opportunities and Risks of Foundation Models, Rishi Bommasani et al. (2021). It conceptualizes foundational risks including emergence, opacity, and homogenization that TrustLLM's trustworthiness principles aim to audit and govern.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). It demonstrates automated semantic jailbreaking techniques that underpin the adversarial robustness evaluations formalized within TrustLLM.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). HarmBench advances TrustLLM's safety and robustness evaluation by introducing a standardized testbed specifically for automated red teaming and dynamic refusal mechanisms.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). This work probes beyond broad trustworthiness benchmarks by designing adaptive, model-specific jailbreak attacks that circumvent contemporary safety alignments.
- Paper: Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection, Zekun Li et al. (2024). This study specifically evaluates instruction-following robustness against prompt injection in retrieved contexts, targeting the retrieval-augmented reliability highlighted in TrustLLM.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). It builds directly on the factuality and hallucination challenges identified in TrustLLM by implementing automated preference optimization pipelines to improve factual reliability.
- Paper: Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures, Dominik Schwarz (2025). This work extends LLM safety auditing into multi-stage and agentic architectures, analyzing systemic cross-stage vulnerabilities beyond single-turn benchmark evaluations.
- Paper: Metacognition in LLMs: Foundations, Progress, and Opportunities, Gabrielle Kaili-May Liu et al. (2026). It provides a deep theoretical and empirical continuation of TrustLLM's findings on model overconfidence and self-knowledge limitations through the lens of metacognition.
