Insights into a radiology-specialised multimodal large language model with sparse autoencoders

Kenza BouzidShruthi BannurFelix MeissenDaniel Coelho de CastroAnton SchwaighoferJavier Alvarez-ValleStephanie Hyland

article2025arXiv2 citations

Applies sparse autoencoders to the radiology multimodal model MAIRA-2 to identify human-interpretable representations of pathologies and medical devices while testing whether these internal features can steer clinical text generation.

Listen

Artificial intelligence systems that generate automated draft radiology reports have the potential to ease clinical workloads and improve hospital efficiency. However, their internal decision-making processes remain opaque, raising serious safety, trust, and compliance concerns in high-stakes clinical settings. Mechanistic interpretability seeks to address this by reverse-engineering model computations into distinct, understandable concepts. The article evaluates whether sparse autoencoders—tools that translate dense internal neural network signals into sparse, human-interpretable features—can successfully isolate clinical concepts within MAIRA-2, a leading multimodal model specialized in chest X-ray reporting, and whether these discovered features can be manipulated to reliably control model outputs.

To conduct this evaluation, the researchers extracted hidden internal representations from the middle layer of MAIRA-2 using over 158,000 paired chest radiographs and clinical reports. They trained a sparse autoencoder to isolate 16,384 distinct latent features and deployed a large language model to automatically generate descriptive labels and interpretability scores for 99.5% of them. The authors then conducted model steering experiments across 7,906 validation studies by modifying internal signals with positive and negative steering vectors to evaluate whether specific clinical concepts could be added, amplified, or suppressed during report generation, using an automated judge to measure intended and unintended textual modifications.

Key findings show that while interpretable concepts exist within the model, they represent a small fraction of the overall internal architecture. Only 1.8% of the discovered features achieved high interpretability scores, while 46% scored at or below random baseline levels. Successfully labeled features captured fine-grained clinical details, such as specific medical devices, anatomical changes, and distinct pathologies like pleural effusions. However, attempting to steer the model using these features produced highly inconsistent results: purely on-target modifications were rare, peaking at only 11.3% of cases for the best-performing feature. Instead, steering attempts frequently generated significant off-target side effects, often fabricating new abnormalities, altering unrelated clinical details, or producing no observable change in approximately 35% of evaluations.

These results demonstrate that direct internal feature intervention is not yet a dependable mechanism for controlling domain-specific clinical models. The high rate of unintended hallucinations and omissions introduces substantial clinical safety risks, differing from the more optimistic steering outcomes often reported in general language domains. The findings imply that internal medical concepts are either entangled across multiple latent dimensions or behave in complex, non-linear ways that resist simple vector adjustments.

Consequently, healthcare organizations and developers should not rely on sparse autoencoder feature steering for operational safety guardrails or clinical output control at this stage. Further technical work is necessary before deployment, including the development of more advanced, multimodal-aware feature explanation methods, refined exemplar sampling, and localized steering interventions that target specific generation steps rather than entire sequences. Confidence in the initial discovery of granular medical features remains high, but substantial caution is advised regarding any claims of reliable output control.

Cover for Insights into a radiology-specialised multimodal large language model with sparse autoencoders

Abstract

Interpretability can improve the safety, transparency and trust of AI models, which is especially important in healthcare applications where decisions often carry significant consequences. Mechanistic interpretability, particularly through the use of sparse autoencoders (SAEs), offers a promising approach for uncovering human-interpretable features within large transformer-based models. In this study, we apply Matryoshka-SAE to the radiology-specialised multimodal large language model, MAIRA-2, to interpret its internal representations. Using large-scale automated interpretability of the SAE features, we identify a range of clinically relevant concepts - including medical devices (e.g., line and tube placements, pacemaker presence), pathologies such as pleural effusion and cardiomegaly, longitudinal changes and textual features. We further examine the influence of these features on model behaviour through steering, demonstrating directional control over generations with mixed success. Our results reveal practical and methodological challenges, yet they offer initial insights into the internal concepts learned by MAIRA-2 - marking a step toward deeper mechanistic understanding and interpretability of a radiology-adapted multimodal large language model, and paving the way for improved model transparency. We release the trained SAEs and interpretations: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Materials and methods
  • 3.1 MAIRA-2 model
  • 3.2 Source dataset
  • 3.3 Extraction and filtering of token representations
  • 3.4 Sparse autoencoder architecture
  • 3.5 SAE training
  • 3.6 Automated interpretation of SAE features
  • 3.7 Steering
  • 3.8 Evaluating steering success
  • 4 Findings
  • 4.1 MAIRA-2 contains interpretable features
  • 4.2 Steering success depends on the feature
  • 5 Discussion and conclusion
  • References
  • A Additional experimental setup details
  • A.1 Representation filtering
  • A.2 SAE Hyperparameters
  • A.3 Auxiliary Loss
  • B Feature statistics
  • C Automated interpretability
  • C.1 Selecting and preparing feature exemplars
  • C.2 Generating interpretations
  • C.3 Scoring interpretations
  • D Automated steering evaluation
  • E Feature steering
  • E.1 Selection of features for steering
  • E.2 Detailed feature steering results

Knowls

  1. Knowl 1 — Matryoshka-SAE Training Objective with Auxiliary Dead-Latent Loss

    equation

    The Matryoshka Sparse Autoencoder (Matryoshka-SAE) trains a set of nested dictionaries M={m1,m2,…,m∣M∣=m}\mathcal{M} = \{m_1, m_2, \dots, m_{|\mathcal{M}|} = m\} of increasing sizes m1<m2<⋯<mm_1 < m_2 < \dots < m simultaneously to represent an input activation vector x∈Rnx \in \mathbb{R}^n. The loss function combines multi-scale reconstruction mean squared error with an auxiliary loss Laux\mathcal{L}_{aux} to revive inactive (dead) latent features:

    L(x)=∑mj∈M∥x−(W1:mjdecf(x)1:mj+bdec)∥22+αLaux\mathcal{L}(x) = \sum_{m_j \in \mathcal{M}} \|x - (W_{1:m_j}^{dec} f(x)_{1:m_j} + b^{dec})\|_2^2 + \alpha \mathcal{L}_{aux}

    where the sparse latent feature representation f(x)∈Rmf(x) \in \mathbb{R}^m is produced by:

    f(x)=σ(Wencx+benc)f(x) = \sigma(W^{enc} x + b^{enc})

    Here, Wenc∈Rm×nW^{enc} \in \mathbb{R}^{m \times n} and benc∈Rmb^{enc} \in \mathbb{R}^m are the encoder weight matrix and bias vector, Wdec∈Rn×mW^{dec} \in \mathbb{R}^{n \times m} and bdec∈Rnb^{dec} \in \mathbb{R}^n are the decoder weight matrix and bias vector, and 1:mj1:m_j denotes slicing the first mjm_j dictionary features. The activation function σ\sigma is the BatchTopK function, which enforces an average sparsity of kk active features per sample across a batch of size bb. The auxiliary loss Laux\mathcal{L}_{aux} measures the reconstruction error of the top-kauxk_{aux} dead features on the model residual:

    Laux=∥e−e^∥22\mathcal{L}_{aux} = \|e - \hat{e}\|_2^2

    where e=x−x^e = x - \hat{x} is the full reconstruction error from the primary dictionary, and e^=Wdeckz\hat{e} = W_{dec}^k z is the reconstruction computed using the top-kauxk_{aux} dead latents. A latent neuron is designated as dead if it has not activated for a predefined threshold of consecutive tokens (e.g., 100,000 tokens).

  2. Knowl 2 — Sparse Autoencoder Setup on MAIRA-2 Residual Stream

    experimental setup

    The base model under study is MAIRA-2, a multimodal large language model for grounded and non-grounded chest X-ray (CXR) report generation. MAIRA-2 couples a radiology-specific vision encoder (producing 1,369 visual tokens per image view) via an MLP adapter to a 32-layer, 7-billion parameter language decoder (hidden dimension n=4096n = 4096) initialized from Vicuna v1.5.

    Dense representations are extracted from the residual stream at the output of the middle layer (Layer 15). All token activation vectors are scaled by a normalization factor of 22.3422.34 (the training set mean ℓ2\ell_2 norm). A Matryoshka-SAE is trained on these normalized representations using the following hyperparameters:

    • Architecture: Matryoshka-SAE with BatchTopK activation
    • Input dimension: n=4096n = 4096
    • Expansion factor: ef=4ef = 4, yielding total latent dictionary size m=16,384m = 16,384
    • Active features (sparsity parameter): k=256k = 256
    • Nested dictionary fractions: [1/2,1/4,1/8,1/16,1/16][1/2, 1/4, 1/8, 1/16, 1/16]
    • Auxiliary loss coefficient: α=0.03125\alpha = 0.03125
    • Training duration: 1 epoch
    • Batch size: 8,192
    • Dead feature threshold: 100,000 tokens
    • Training corpus: MIMIC-CXR (158,555 training studies and 7,906 validation studies).
  3. Knowl 3 — Multimodal Token Representation Filtering for Sparse Autoencoders

    model/method

    To train sparse autoencoders on instruction-tuned multimodal models without learning superficial artifacts, input token sequences are filtered prior to dictionary learning. In the MAIRA-2 prompt pipeline, raw sequences reach up to 5,099 tokens (averaging 3,358±8693,358 \pm 869 tokens), comprising up to three CXR image views (1,369 tokens per view), clinical text sections (indication, technique, comparison), and generated findings.

    The filtering protocol discards:

    1. System prompt text and task instructions.
    2. Beginning-of-sequence, end-of-sequence, and chat template delimiters.
    3. Intermediate visual tokens: out of each 1,369-token image block, only the final image token is retained to ensure the SAE processes complete visual representations rather than partial image patches.

    This filtering reduces the token count per sample from thousands to an average of 176 tokens, yielding 34.7 million token representations for SAE training and 1.7 million token representations for validation on MIMIC-CXR.

  4. Knowl 4 — Automated Feature Interpretation and Detection-Scoring Protocol

    algorithm

    To assign human-understandable labels and measure feature interpretability at scale, a two-stage automated LLM pipeline (using text-only GPT-4o) is employed:

    Input: Trained SAE latent feature fif_i, dataset of N=500,000N = 500,000 token activations
    Output: Natural language explanation EiE_i, detection interpretability score F1,iF_{1,i}
    // Phase 1: Feature Interpretation Generation
    Sactive←S_{active} \leftarrow sample 25 token contexts from top decile of activation distribution of fif_i
    Szero←S_{zero} \leftarrow sample 25 token contexts where fi(x)=0f_i(x) = 0
    Sexemplar←Sactive∪SzeroS_{exemplar} \leftarrow S_{active} \cup S_{zero}
    for each sample s∈Sexemplars \in S_{exemplar} do
        Replace all visual tokens in sequence with literal string '<image>'
        Truncate context to full preceding sequence plus 100 characters after target token
        Highlight target token with double brackets '[[token]]'
        Scale continuous activation value fi(s)f_i(s) to integer score ∈[0,9]\in [0, 9]
    Prompt interpretation LLM with SexemplarS_{exemplar} and few-shot clinical examples
    Ei←E_i \leftarrow LLM-generated concise explanation and rationale
    // Phase 2: Detection-Based Quality Scoring
    Tpos←T_{pos} \leftarrow sample 100 non-overlapping token contexts from top quintile of fif_i
    Tneg←T_{neg} \leftarrow sample 100 non-overlapping token contexts where fi(x)=0f_i(x) = 0
    Teval←Tpos∪TnegT_{eval} \leftarrow T_{pos} \cup T_{neg}
    for each test sample t∈Tevalt \in T_{eval} do
        Prompt scoring LLM with explanation EiE_i and sample tt to classify y^∈{0,1}\hat{y} \in \{0, 1\}
    F1,i←F_{1,i} \leftarrow Compute detection F1F_1 score between ground truth binary activations and predictions y^\hat{y}
    return EiE_i, F1,iF_{1,i}
  5. Knowl 5 — Interpretability Distribution of Learned SAE Features in MAIRA-2

    empirical result

    Automated interpretation and detection-scoring across 16,299 evaluated SAE features (out of 16,384 total features) in MAIRA-2 revealed that high interpretability is sparse:

    • Only 288 features (1.8%1.8\%) achieved a detection F1F_1 score above 0.750.75.
    • 7,500 features (46.0%46.0\%) scored an F1F_1 below 0.500.50, corresponding to chance performance on the balanced binary detection evaluation set.
    • Scoring LLM predictions exhibited consistently higher recall than precision, reflecting a tendency for generated feature explanations to be overly broad or non-specific.
    • Highly interpretable features (F1>0.85F_1 > 0.85) mapped to specific anatomical and pathological findings (e.g., f1336f_{1336} for aortic tortuosity/calcification with F1=0.89F_1 = 0.89; f14586f_{14586} for plate-like atelectasis with F1=0.94F_1 = 0.94), medical line and tube placements (e.g., f11240f_{11240} for chest tube placement/removal with F1=0.94F_1 = 0.94; f8757f_{8757} for central line placement in mid-SVC with F1=0.92F_1 = 0.92), temporal comparisons, and linguistic discourse markers (e.g., f12106f_{12106} for use of 'however' indicating suspicious findings with F1=0.93F_1 = 0.93).
    • Features with low F1F_1 scores predominantly received generic descriptions of comparative radiology reporting or task prompt boilerplate.
  6. Knowl 6 — Residual Stream Feature Steering and LLM-as-a-Judge Evaluation

    model/method

    To steer MAIRA-2 generations toward or away from a target concept represented by SAE latent feature fif_i, the corresponding column vector Widec∈RnW^{dec}_i \in \mathbb{R}^n from the SAE decoder matrix is utilized as a steering vector. During autoregressive decoding, the modified hidden state ht′h_t' at layer 15 for each token position tt is computed as:

    ht′=ht+αWidech_t' = h_t + \alpha W^{dec}_i

    where α∈R\alpha \in \mathbb{R} is the steering coefficient. A value of α=+10\alpha = +10 is used for positive steering (concept induction) and α=−10\alpha = -10 for negative steering (concept suppression).

    Evaluation is performed using an automated LLM judge (GPT-4o) provided with the target feature concept, the original unsteered report, and the steered report. The judge assigns two continuous scores from 0.00.0 to 1.01.0:

    1. On-target score: the extent to which the modification specifically promotes (for α>0\alpha > 0) or suppresses (for α<0\alpha < 0) the targeted concept.
    2. Off-target score: the degree of significant unrelated content alteration, hallucination, or omission in the remaining report.

    Scores are binarized at a threshold of 0.10.1 to classify outcomes into four disjoint categories: (1) only on-target changes, (2) both on- and off-target changes, (3) only off-target changes, and (4) no changes.

  7. Knowl 7 — Prevalence of Off-Target Side Effects in SAE-Based Multimodal Steering

    empirical result

    Evaluation of feature steering (α=+10\alpha = +10) across 67 interpretable features on 7,906 validation studies revealed that steering frequently produces unintended side effects:

    • Pure on-target modification without collateral changes was rare, reaching a maximum rate of 11.3%11.3\% (observed for f10709f_{10709}, clear lungs finding).
    • In approximately 35%35\% of all evaluations across features, steering produced no observable modification to the generated report.
    • Combined on- and off-target changes occurred frequently (e.g., 27.4%27.4\% in f6412f_{6412}, pleural effusion detection).
    • For several features (including f11509f_{11509} for rib fractures, f10643f_{10643} for telephone notifications, and f13506f_{13506} for scoliosis), more than 50%50\% of steered generations produced exclusively off-target modifications (such as confabulating pneumothorax, effusion, or surgical hardware) with zero on-target change.
  8. Knowl 8 — Bidirectional Correlation and Activation Frequency Effects in Feature Steering

    empirical result

    Statistical analysis of steering dynamics in MAIRA-2 revealed two consistent operational characteristics:

    1. Bidirectional Symmetry: On-target steering efficacy under positive intervention (α=+10\alpha = +10) strongly correlates with on-target suppression efficacy under negative intervention (α=−10\alpha = -10) across evaluated features (Spearman's ρ=0.90,p≤0.05,n=67\rho = 0.90, p \le 0.05, n = 67). A similar but weaker correlation holds for off-target effects (Spearman's ρ=0.71,p≤0.05,n=67\rho = 0.71, p \le 0.05, n = 67). This indicates that the directional vectors retain consistent linear properties in latent space across positive and negative shifts.
    2. Frequency Dependence: For features with high interpretability (F1>0.85F_1 > 0.85), the frequency with which a feature activates in the training dataset is significantly positively correlated with its on-target steering score (Spearman's ρ=0.40,p≤0.05,n=50\rho = 0.40, p \le 0.05, n = 50, assessed via 9,999 permutation tests). Feature activation frequency did not exhibit a significant correlation with off-target side-effect scores.
  9. Knowl 9 — Discovered Clinical Feature Taxonomy and Steering Characteristics

    data/table

    The following table details representative interpretable SAE features extracted from Layer 15 of MAIRA-2, grouped by clinical domain, showing validation detection F1F_1, active training sample counts (out of 500,000 sampled tokens), natural language interpretations, and steering behaviors:

    Category Feature ID Detection F1F_1 # Active Automated Explanation
    Medical Devices f12062f_{12062} 0.95 1,028 Presence or repositioning of pigtail catheters in chest imaging.
    Medical Devices f11240f_{11240} 0.94 1,365 Descriptions of findings related to chest tube placement or removal.
    Medical Devices f8757f_{8757} 0.92 420 Central line or catheter placement described with specific position (e.g. 'mid SVC').
    Medical Devices f3246f_{3246} 0.80 27,373 Evaluation of pacemaker or ICD lead positions in chest X-rays.
    Abnormal Findings f13515f_{13515} 0.97 159 Elevation of the hemidiaphragm.
    Abnormal Findings f14586f_{14586} 0.94 533 Presence of plate-like or linear atelectasis.
    Abnormal Findings f12585f_{12585} 0.93 322 Unfolded or tortuous thoracic aorta in radiology reports.
    Abnormal Findings f11509f_{11509} 0.89 1,123 Observations of rib fractures in chest imaging reports.
    Abnormal Findings f6108f_{6108} 0.86 11,089 Findings of pulmonary vascular congestion or pulmonary edema.
    Abnormal Findings f4875f_{4875} 0.86 4,444 Cardiomegaly or enlarged cardiac silhouette.
    Abnormal Findings f6412f_{6412} 0.84 5,922 Detection of pleural effusions on imaging studies.
    Normal Findings f11891f_{11891} 0.87 1,374 Normal imaging findings with emphasis on absence of acute pathology.
    Normal Findings f10709f_{10709} 0.87 1,761 Findings indicate clear lungs with no signs of pleural effusion, pneumothorax, consolidation, or pulmonary edema.
    Temporal Changes f1646f_{1646} 0.79 30,532 Interval change in disease findings from prior imaging.
    Temporal Changes f1599f_{1599} 0.79 98,759 Describing findings without comparison to prior images.
    Textual Features f12106f_{12106} 0.88 660 Use of 'however' in clinical findings indicating possible issues needing further investigation.
    Textual Features f9473f_{9473} 0.86 677 Use of 'possible' or 'possibly' indicating uncertainty.

    These results show that SAE latents capture specific radiological concepts (such as anatomical device positions or isolated pathologies) with higher granularity than standard broad classification labels (e.g., CheXpert), while also capturing temporal reasoning and linguistic hedge tokens.

  10. Knowl 10 — Impact of Image Findings Prefixing on Automated Feature Interpretation

    empirical result

    To evaluate whether providing visual context to a text-only interpretation LLM improves concept discovery for a multimodal model, experiments compared two exemplar presentation strategies:

    1. Standard text-only formatting, where visual token embeddings are replaced by the string identifier <image>.
    2. Image description augmentation, where the ground-truth radiological 'Findings' section of the report (serving as an approximate description of image content) was prefixed to the data sample.

    Providing the prefixed image findings description yielded no improvement in the number of interpretable SAE features discovered across any detection F1F_1 threshold. This occurs in part because the SAE is trained and evaluated on both the prompt and target tokens of MAIRA-2, meaning that in many training samples the surrounding textual sequence already incorporates relevant descriptions of visual findings.

Coverage note — None was omitted; all key contributions—including the Matryoshka-SAE formulation on MAIRA-2, token filtering, automated interpretation and detection scoring, interpretability statistics, steering mechanics, LLM-as-a-judge evaluation, steering failure modes, and statistical correlations—are fully captured.

References

  1. 1.Abdulaal, A., Fry, H., Montaña-Brown, N., Ijishakin, A., Gao, J., Hyland, S., Alexander, D. C., and Castro, D. C. An X-ray is worth 15 features: Sparse autoencoders for interpretable radiology report generation. arXiv preprint arXiv:2410.03334, 2024.
  2. 2.Adams, E., Bai, L., Lee, M., Yu, Y., and AlQuraishi, M. From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language Models, February 2025. URL https://www.biorxiv.org/content/10.1101/2025.02.06.636901v1. Pages: 2025.02.06.636901 Section: New Results.
  3. 3.Anders, E., Neo, C., Hoelscher-Obermaier, J., and Howard, J. N. Sparse autoencoders find composed features in small toy models, 2024. URL https://www.lesswrong.com/posts/a5wwqza2cY3W7L9cj.
  4. 4.Bannur, S., Bouzid, K., Castro, D. C., Schwaighofer, A., Bond-Taylor, S., Ilse, M., Pérez-García, F., Salvatelli, V., Sharma, H., Meissen, F., Ranjit, M., Srivastav, S., Gong, J., Falck, F., Oktay, O., Thieme, A., Lungren, M. P., Wetscherek, M. T., Alvarez-Valle, J., and Hyland, S. L. MAIRA-2: Grounded Radiology Report Generation, June 2024. URL http://arxiv.org/abs/2406.04449.
  5. 5.Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W. Language models can explain neurons in language models. https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html, 2023.
  6. 6.Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread, 2023. https://transformercircuits.pub/2023/monosemantic-features/index.html.
  7. 7.Bussmann, B., Leask, P., and Nanda, N. BatchTopK sparse autoencoders. arXiv preprint arXiv:2412.06410, 2024.
  8. 8.Bussmann, B., Nabeshima, N., Karvonen, A., and Nanda, N. Learning Multi-Level Features with Matryoshka Sparse Autoencoders, March 2025. URL http://arxiv.org/abs/2503.17547. arXiv:2503.17547 [cs].
  9. 9.Chanin, D., Wilken-Smith, J., Dulka, T., Bhatnagar, H., and Bloom, J. I. A is for absorption: Studying feature splitting and absorption in sparse autoencoders. In NeurIPS 2024 Workshop on Interpretable AI, December 2024. URL https://openreview.net/forum?id=Wzav8fesTL.
  10. 10.Chen, Z., Varma, M., Delbrouck, J.-B., Paschali, M., Blankemeier, L., Van Veen, D., Valanarasu, J. M. J., Youssef, A., Cohen, J. P., Reis, E. P., et al. CheXagent: Towards a foundation model for chest x-ray interpretation. arXiv preprint arXiv:2401.12208, 2024.
  11. 11.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality, March 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/.
  12. 12.Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L. Sparse autoencoders find highly interpretable features in language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=F76bwRSLeK.
  13. 13.Durmus, E., Tamkin, A., Clark, J., Wei, J., Marcus, J., Batson, J., Handa, K., Lovitt, L., Tong, M., McCain, M., Rausch, O., Huang, S., Bowman, S., Ritchie, S., Henighan, T., and Ganguli, D. Evaluating feature steering: A case study in mitigating social biases, 2024. URL https://anthropic.com/research/evaluating-feature-steering.
  14. 14.Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021.
  15. 15.Fiotto-Kaufman, J., Loftus, A. R., Todd, E., Brinkmann, J., Juang, C., Pal, K., Rager, C., Mueller, A., Marks, S., Sharma, A. S., et al. NNsight and NDIF: Democratizing access to foundation model internals. arXiv preprint arXiv:2407.14561, 2024.
  16. 16.Gao, L., la Tour, T. D., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J. Scaling and evaluating sparse autoencoders, June 2024. URL http://arxiv.org/abs/2406.04093. arXiv:2406.04093 [cs].
  17. 17.Gur-Arieh, Y., Mayan, R., Agassy, C., Geiger, A., and Geva, M. Enhancing automated interpretability with output-centric feature descriptions. arXiv preprint arXiv:2501.08319, 2025.
  18. 18.He, Z., Shu, W., Ge, X., Chen, L., Wang, J., Zhou, Y., Liu, F., Guo, Q., Huang, X., Wu, Z., Jiang, Y.-G., and Qiu, X. Llama Scope: Extracting millions of features from Llama-3.1-8B with sparse autoencoders, 2024. URL https://arxiv.org/abs/2410.20526.
  19. 19.Hyland, S. L., Bannur, S., Bouzid, K., Castro, D. C., Ranjit, M., Schwaighofer, A., Pérez-García, F., Salvatelli, V., Srivastav, S., Thieme, A., et al. MAIRA-1: A specialised large multimodal model for radiology report generation. arXiv:2311.13668, 2023. URL https://arxiv.org/abs/2311.13668.
  20. 20.Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., et al. CheXpert: A large chest radiograph dataset with uncertainty labels and expert comparison. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pp. 590–597, 2019.
  21. 21.Jiang, Y., Chen, C., Nguyen, D., Mervak, B. M., and Tan, C. GPT-4V cannot generate radiology reports yet. ArXiv, abs/2407.12176, 2024. URL https://api.semanticscholar.org/CorpusID:271244474.
  22. 22.Johnson, A. E. W., Pollard, T. J., Berkowitz, S. J., Mark, R. G., and Horng, S. MIMIC-CXR database (version 2.0.0). PhysioNet, 2019.
  23. 23.Lad, V., Gurnee, W., and Tegmark, M. The remarkable robustness of LLMs: Stages of inference? arXiv preprint arXiv:2406.19384, 2024.
  24. 24.Le, N. M., Patel, N., Shen, C., Martin, B., Eng, A., Shah, C., Grullon, S., and Juyal, D. Learning biologically relevant features in a pathology foundation model using sparse autoencoders. In NeurIPS 2024 Workshop on Advancements In Medical Foundation Models, December 2024. URL https://openreview.net/forum?id=daV16mhUBd.
  25. 25.Lee, H., Battle, A., Raina, R., and Ng, A. Efficient sparse coding algorithms. Advances in neural information processing systems, 19, 2006.
  26. 26.Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36:41451–41530, 2023a.
  27. 27.Li, M., Lin, B., Chen, Z., Lin, H., Liang, X., and Chang, X. Dynamic graph enhanced contrastive learning for chest x-ray report generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3334–3343, 2023b.
  28. 28.Lieberum, T., Rajamanoharan, S., Conmy, A., Smith, L., Sonnerat, N., Varma, V., Kramár, J., Dragan, A., Shah, R., and Nanda, N. Gemma Scope: Open sparse autoencoders everywhere all at once on Gemma 2. arXiv preprint arXiv:2408.05147, 2024.
  29. 29.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. In Advances in Neural Information Processing Systems, volume 36, pp. 34892–34916, 2023.
  30. 30.Lou, H., Li, C., Ji, J., and Yang, Y. SAE-V: Interpreting Multimodal Models for Enhanced Alignment, February 2025. URL http://arxiv.org/abs/2502.17514. arXiv:2502.17514 [cs].
  31. 31.Marks, S., Karvonen, A., and Mueller, A. dictionary_learning. https://github.com/saprmarks/dictionary_learning, 2024.
  32. 32.Minder, J., Dumas, C., Juang, C., Chugtai, B., and Nanda, N. Robustly identifying concepts introduced during chat fine-tuning using crosscoders. arXiv preprint arXiv:2504.02922, 2025.
  33. 33.Mudide, A., Engels, J., Michaud, E. J., Tegmark, M., and de Witt, C. S. Efficient dictionary learning with switch sparse autoencoders. arXiv preprint arXiv:2410.08201, 2024.
  34. 34.O’Brien, K., Majercak, D., Fernandes, X., Edgar, R., Chen, J., Nori, H., Carignan, D., Horvitz, E., and Poursabzi-Sangde, F. Steering language model refusal with sparse autoencoders. arXiv preprint arXiv:2411.11296, 2024.
  35. 35.Pach, M., Karthik, S., Bouniot, Q., Belongie, S., and Akata, Z. Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models, April 2025. URL http://arxiv.org/abs/2504.02821. arXiv:2504.02821 [cs].
  36. 36.Parekh, J., Khayatan, P., Shukor, M., Newson, A., and Cord, M. A concept-based explainability framework for large multimodal models. Advances in Neural Information Processing Systems, 37:135783–135818, 2024.
  37. 37.Park, K., Choe, Y. J., and Veitch, V. The linear representation hypothesis and the geometry of large language models. In Proceedings of the 41st International Conference on Machine Learning, pp. 39643–39666. PMLR, July 2024. URL https://proceedings.mlr.press/v235/park24c.html.
  38. 38.Paulo, G., Mallen, A., Juang, C., and Belrose, N. Automatically Interpreting Millions of Features in Large Language Models, December 2024. URL http://arxiv.org/abs/2410.13928. arXiv:2410.13928 [cs].
  39. 39.Pérez-García, F., Sharma, H., Bond-Taylor, S., Bouzid, K., Salvatelli, V., Ilse, M., Bannur, S., Castro, D. C., Schwaighofer, A., Lungren, M. P., et al. Exploring scalable medical image encoders beyond text supervision. Nature Machine Intelligence, 7:119–130, 2025. doi: 10.1038/s42256-024-00965-w.
  40. 40.Quinn, T. P., Senadeera, M., Jacobs, S., Coghlan, S., and Le, V. Trust and medical ai: the challenges we face and the expertise needed to overcome them. Journal of the American Medical Informatics Association, 28(4):890–894, 2021.
  41. 41.Shaham, T. R., Schwettmann, S., Wang, F., Rajaram, A., Hernandez, E., Andreas, J., and Torralba, A. A multimodal automated interpretability agent. In Forty-first International Conference on Machine Learning, 2024.
  42. 42.Simon, E. and Zou, J. InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders, November 2024. URL https://www.biorxiv.org/content/10.1101/2024.11.14.623630v1. Pages: 2024.11.14.623630 Section: New Results.
  43. 43.Stevens, S., Chao, W.-L., Berger-Wolf, T., and Su, Y. Sparse autoencoders for scientifically rigorous interpretation of vision models. arXiv preprint arXiv:2502.06755, 2025.
  44. 44.Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T. Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet. Transformer Circuits Thread, 2024. URL https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html.
  45. 45.Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P.-C., Carroll, A., Lau, C., Tanno, R., Ktena, I., Palepu, A., Mustafa, B., Chowdhery, A., Liu, Y., Kornblith, S., Fleet, D., Mansfield, P., Prakash, S., Wong, R., Virmani, S., et al. Towards generalist biomedical AI. NEJM AI, 1(3):AIoa2300138, February 2024. doi: 10.1056/AIoa2300138.
  46. 46.Turner, A. M., Thiergart, L., Leech, G., Udell, D., Vazquez, J. J., Mini, U., and MacDiarmid, M. Steering language models with activation engineering. arXiv preprint arXiv:2308.10248, 2023.
  47. 47.Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. arXiv preprint arXiv:2211.00593, 2022.
  48. 48.Wang, Z., Liu, L., Wang, L., and Zhou, L. Metransformer: Radiology report generation by transformer with multiple learnable expert tokens. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11558–11567, 2023.
  49. 49.Wattenberg, M. and Viégas, F. B. Relational composition in neural networks: A survey and call to action, 2024. URL https://arxiv.org/abs/2407.14662.
  50. 50.Wu, Z., Arora, A., Geiger, A., Wang, Z., Huang, J., Jurafsky, D., Manning, C. D., and Potts, C. AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders, March 2025. URL http://arxiv.org/abs/2501.17148. arXiv:2501.17148 [cs].
  51. 51.Yan, Q., He, X., Yue, X., and Wang, X. E. Worse than random? an embarrassingly simple probing evaluation of large multimodal models in medical VQA. ArXiv, abs/2405.20421, 2024. URL https://api.semanticscholar.org/CorpusID:270199350.
  52. 52.Yang, L., Xu, S., Sellergren, A., Kohlberger, T., Zhou, Y., Ktena, I., Kiraly, A., Ahmed, F., Hormozdiari, F., Jaroensri, T., et al. Advancing multimodal medical capabilities of Gemini. arXiv preprint arXiv:2405.03162, 2024.
  53. 53.Yildirim, N., Richardson, H., Wetscherek, M. T., Bajwa, J., Jacob, J., Pinnock, M. A., Harris, S., Coelho De Castro, D., Bannur, S., Hyland, S., et al. Multimodal healthcare AI: Identifying and designing clinically relevant vision-language applications for radiology. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, pp. 1–22, 2024.
  54. 54.Zhang, K., Shen, Y., Li, B., and Liu, Z. Large Multimodal Models Can Interpret Features in Large Multimodal Models, November 2024. URL http://arxiv.org/abs/2411.14982. arXiv:2411.14982 [cs].
  55. 55.Zhou, H.-Y., Adithan, S., Acosta, J. N., Topol, E. J., and Rajpurkar, P. A generalist learner for multifaceted medical image interpretation. arXiv preprint arXiv:2405.07988, 2024.

Citation

MLA
Bouzid, K., et al. “Insights into a Radiology-specialised Multimodal Large Language Model with Sparse Autoencoders”. arXiv, 2025, http://arxiv.org/abs/2507.12950v2.
APA
Bouzid, K., Bannur, S., Meissen, F., Castro, D. C. de ., Schwaighofer, A., Alvarez-Valle, J., & Hyland, S. L. (2025). Insights into a radiology-specialised multimodal large language model with sparse autoencoders. arXiv. http://arxiv.org/abs/2507.12950v2
Chicago
Bouzid, K., S. Bannur, F. Meissen, et al. 2025. “Insights into a Radiology-specialised Multimodal Large Language Model with Sparse Autoencoders”. arXiv. http://arxiv.org/abs/2507.12950v2.
Harvard
Bouzid, K. et al. (2025) “Insights into a radiology-specialised multimodal large language model with sparse autoencoders”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2507.12950v2.
Vancouver
1. Bouzid K, Bannur S, Meissen F, Castro DC de, Schwaighofer A, Alvarez-Valle J, Hyland SL (2025) Insights into a radiology-specialised multimodal large language model with sparse autoencoders. arXiv

BibTeX

@article{bouzid2025insights,
  title = {Insights into a radiology-specialised multimodal large language model with sparse autoencoders},
  author = {Bouzid, Kenza and Bannur, Shruthi and Meissen, Felix and Castro, Daniel Coelho de and Schwaighofer, Anton and Alvarez-Valle, Javier and Hyland, Stephanie L.},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2507.12950v2},
  eprint = {2507.12950}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission