Insights into a radiology-specialised multimodal large language model with sparse autoencoders
Kenza BouzidShruthi BannurFelix MeissenDaniel Coelho de CastroAnton SchwaighoferJavier Alvarez-ValleStephanie Hyland
Applies sparse autoencoders to the radiology multimodal model MAIRA-2 to identify human-interpretable representations of pathologies and medical devices while testing whether these internal features can steer clinical text generation.
Artificial intelligence systems that generate automated draft radiology reports have the potential to ease clinical workloads and improve hospital efficiency. However, their internal decision-making processes remain opaque, raising serious safety, trust, and compliance concerns in high-stakes clinical settings. Mechanistic interpretability seeks to address this by reverse-engineering model computations into distinct, understandable concepts. The article evaluates whether sparse autoencoders—tools that translate dense internal neural network signals into sparse, human-interpretable features—can successfully isolate clinical concepts within MAIRA-2, a leading multimodal model specialized in chest X-ray reporting, and whether these discovered features can be manipulated to reliably control model outputs.
To conduct this evaluation, the researchers extracted hidden internal representations from the middle layer of MAIRA-2 using over 158,000 paired chest radiographs and clinical reports. They trained a sparse autoencoder to isolate 16,384 distinct latent features and deployed a large language model to automatically generate descriptive labels and interpretability scores for 99.5% of them. The authors then conducted model steering experiments across 7,906 validation studies by modifying internal signals with positive and negative steering vectors to evaluate whether specific clinical concepts could be added, amplified, or suppressed during report generation, using an automated judge to measure intended and unintended textual modifications.
Key findings show that while interpretable concepts exist within the model, they represent a small fraction of the overall internal architecture. Only 1.8% of the discovered features achieved high interpretability scores, while 46% scored at or below random baseline levels. Successfully labeled features captured fine-grained clinical details, such as specific medical devices, anatomical changes, and distinct pathologies like pleural effusions. However, attempting to steer the model using these features produced highly inconsistent results: purely on-target modifications were rare, peaking at only 11.3% of cases for the best-performing feature. Instead, steering attempts frequently generated significant off-target side effects, often fabricating new abnormalities, altering unrelated clinical details, or producing no observable change in approximately 35% of evaluations.
These results demonstrate that direct internal feature intervention is not yet a dependable mechanism for controlling domain-specific clinical models. The high rate of unintended hallucinations and omissions introduces substantial clinical safety risks, differing from the more optimistic steering outcomes often reported in general language domains. The findings imply that internal medical concepts are either entangled across multiple latent dimensions or behave in complex, non-linear ways that resist simple vector adjustments.
Consequently, healthcare organizations and developers should not rely on sparse autoencoder feature steering for operational safety guardrails or clinical output control at this stage. Further technical work is necessary before deployment, including the development of more advanced, multimodal-aware feature explanation methods, refined exemplar sampling, and localized steering interventions that target specific generation steps rather than entire sequences. Confidence in the initial discovery of granular medical features remains high, but substantial caution is advised regarding any claims of reliable output control.
- Paper: Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham et al. (2023). It introduces the foundational framework of training sparse autoencoders on language model activations to extract monosemantic, human-interpretable features and steer model behaviors.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). It establishes the core methodology of representation engineering and activation steering to read internal cognitive concepts and control downstream generation in large models.
- Paper: The Linear Representation Hypothesis and the Geometry of Large Language Models, Kiho Park et al. (2024). It formalizes the linear representation hypothesis underlying mechanistic interpretability and activation manipulation in large language models.
- Paper: A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance Estimation, Thomas Fel et al. (2023). It provides a unified theoretical foundation for unsupervised concept extraction via dictionary learning and concept importance attribution in neural networks.
- Paper: RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations, Jing Huang et al. (2024). It offers a critical diagnostic benchmark for evaluating how well interpretability methods like sparse autoencoders isolate and disentangle internal representations.
- Paper: LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day, Chunyuan Li et al. (2023). It outlines the multimodal architectural and training foundations of biomedical large vision-language assistants that specialize in radiological and clinical reasoning.
- Paper: Scaling and evaluating sparse autoencoders, Leo Gao et al. (2025). It develops advanced scaling recipes, TopK activations, and evaluation metrics for training large-scale sparse autoencoders across broader language models.
- Paper: VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge, Vishwesh Nath et al. (2025). It extends specialized multimodal medical vision-language architectures by integrating clinical expert tools and conversational reasoning across diverse healthcare tasks.
- Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). It critiques the statistical fragility and non-identifiability of modern interpretability methods, including sparse autoencoders, proposing rigorous causal guardrails.
