RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

Jiduan LiuJiahao LiuQifan WangJingang WangWei WuYunsen XianDongyan ZhaoKai ChenRui Yan

article2023ACL54 citations

Proposes RankCSE, an unsupervised sentence representation framework that integrates ranking consistency and listwise ranking distillation into contrastive learning to capture fine-grained semantic similarities and outperform existing baselines on semantic textual similarity tasks.

Listen

Modern natural language processing systems often rely on unsupervised sentence representations to convert text into numerical vectors for tasks such as information retrieval, text clustering, and semantic search. While recent contrastive learning techniques have advanced this area by pulling similar sentences together and pushing dissimilar ones apart, they treat non-matching sentences equally as generic negative samples. This binary distinction prevents models from learning fine-grained ranking relationships—such as distinguishing highly relevant sentences from moderately relevant ones—which limits search precision and recommendation quality.

The article introduces and evaluates RankCSE, a framework designed to learn semantically discriminative sentence embeddings by integrating listwise ranking principles directly into unsupervised contrastive learning. The main objective was to demonstrate that incorporating ranking consistency and ranking distillation allows models to generalize nuanced, fine-grained semantic hierarchies without requiring expensive human-annotated data.

The researchers trained their models using one million unlabeled sentences randomly sampled from English Wikipedia across standard pre-trained architectures, including BERT and RoBERTa. The approach combines standard contrastive loss with two new components: a ranking consistency loss that enforces stable similarity orderings across different network augmentations, and a ranking distillation loss that transfers listwise ranking signals from teacher models. The evaluation spanned seven semantic textual similarity benchmarks and seven downstream transfer classification tasks, measuring performance against established baselines using correlation and accuracy metrics.

The key findings show consistent performance advantages across all benchmarks. First, RankCSE outperformed existing unsupervised baselines across all evaluated architectures, achieving average semantic similarity scores of 80.05% to 80.60% on BERT and 79.73% to 80.60% on RoBERTa. Second, the base version of RankCSE surpassed the much larger baseline model (SimCSE-BERTlarge) on semantic similarity tasks by roughly 2%, illustrating substantial representational efficiency. Third, downstream classification tasks improved to an average accuracy of 87.33% on standard benchmarks. Finally, geometric analysis confirmed that RankCSE establishes a superior balance between bringing related sentences close together and maintaining a uniform spread across the representation space, leading to more robust and stable embeddings.

These results demonstrate that incorporating learning-to-rank methods into unsupervised embedding models enables significant quality gains without labeled training data. For organizations deploying search, matching, or recommendation pipelines, this framework delivers higher precision and relevance while avoiding data labeling costs. Moreover, because smaller base models equipped with RankCSE can outperform standard large models, organizations can lower inference latency and cloud compute expenditures.

Engineering and research teams seeking to improve search or text-similarity pipelines should evaluate RankCSE-style ranking objectives as drop-in upgrades over standard contrastive learning methods. Multi-teacher distillation configurations—such as combining predictions from complementary models—are recommended for optimal performance. Before full production deployment, teams should conduct pilot tests to evaluate the added training-time cost, as calculating teacher ranking signals increases training duration to roughly 2 to 3.7 hours on an A100 GPU compared to simpler single-stage contrastive approaches. Future work should focus on automating teacher model selection and extending listwise ranking objectives to domain-specific corpora.

arXiv: 2305.16726
Cover for RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

Abstract

Unsupervised sentence representation learning is one of the fundamental problems in natural language processing with various downstream applications. Recently, contrastive learning has been widely adopted which derives high-quality sentence representations by pulling similar semantics closer and pushing dissimilar ones away. However, these methods fail to capture the fine-grained ranking information among the sentences, where each sentence is only treated as either positive or negative. In many real-world scenarios, one needs to distinguish and rank the sentences based on their similarities to a query sentence, e.g., very relevant, moderate relevant, less relevant, irrelevant, etc. In this paper, we propose a novel approach, RankCSE, for unsupervised sentence representation learning, which incorporates ranking consistency and ranking distillation with contrastive learning into a unified framework. In particular, we learn semantically discriminative sentence representations by simultaneously ensuring ranking consistency between two representations with different dropout masks, and distilling listwise ranking knowledge from the teacher. An extensive set of experiments are conducted on both semantic textual similarity (STS) and transfer (TR) tasks. Experimental results demonstrate the superior performance of our approach over several state-of-the-art baselines.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminary
  • 4 Methodology
  • 4.1 Problem Formulation
  • 4.2 Contrastive Learning
  • 4.3 Ranking Consistency
  • 4.4 Ranking Distillation
  • 5 Experiment
  • 5.1 Setup
  • 5.2 Main Results
  • 5.3 Analysis and Discussion
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Training Details
  • B DiffCSE Settings for Transfer Tasks
  • C Data Statistics
  • D Training Efficiency
  • E Cosine Similarity Distribution
  • F Case Study
  • G Ranking Tasks
  • H Alignment and Uniformity

Knowls

  1. Knowl 1 — RankCSE Framework and Unified Training Objective

    model/method

    RankCSE is an unsupervised sentence representation learning framework that extends contrastive learning by incorporating ranking consistency and listwise ranking distillation. Given a mini-batch of NN input sentences {xi}i=1N\{x_i\}_{i=1}^N, each sentence xix_i is passed through an encoder f(⋅)f(\cdot) twice with independently sampled dropout masks to produce two representations f(xi)f(x_i) and f(xi)′f(x_i)'.

    The overall optimization objective Lfinal\mathcal{L}_{\text{final}} unifies three distinct loss components: Lfinal=LinfoNCE+βLconsistency+γLrank\mathcal{L}_{\text{final}} = \mathcal{L}_{\text{infoNCE}} + \beta \mathcal{L}_{\text{consistency}} + \gamma \mathcal{L}_{\text{rank}} where β\beta and γ\gamma are balancing hyperparameters (both set to 11 in standard configurations).

    The contrastive loss LinfoNCE\mathcal{L}_{\text{infoNCE}} maximizes the similarity of positive dropout pairs against in-batch negatives: LinfoNCE=−∑i=1Nlog⁡eϕ(f(xi),f(xi)′)/τ1∑j=1Neϕ(f(xi),f(xj)′)/τ1\mathcal{L}_{\text{infoNCE}} = -\sum_{i=1}^N \log \frac{e^{\phi(f(x_i), f(x_i)') / \tau_1}}{\sum_{j=1}^N e^{\phi(f(x_i), f(x_j)') / \tau_1}} where ϕ(u,v)=u⊤v∥u∥∥v∥\phi(u, v) = \frac{u^\top v}{\|u\| \|v\|} is the cosine similarity and τ1\tau_1 is a temperature hyperparameter (default τ1=0.05\tau_1 = 0.05). While LinfoNCE\mathcal{L}_{\text{infoNCE}} pushes all negatives away equally, Lconsistency\mathcal{L}_{\text{consistency}} and Lrank\mathcal{L}_{\text{rank}} preserve fine-grained semantic ranking relations among negatives.

  2. Knowl 2 — Ranking Consistency Objective via Jensen-Shannon Divergence

    equation

    To capture fine-grained semantic similarity distinctions among in-batch negatives without supervised labels, RankCSE enforces ranking consistency between the two similarity lists produced by two stochastic dropout passes of an anchor sentence xix_i. Let the two similarity lists across the mini-batch of size NN be: S(xi)={ϕ(f(xi),f(xj)′)}j=1NS(x_i) = \{\phi(f(x_i), f(x_j)')\}_{j=1}^N S(xi)′={ϕ(f(xi)′,f(xj))}j=1NS(x_i)' = \{\phi(f(x_i)', f(x_j))\}_{j=1}^N where ϕ(⋅,⋅)\phi(\cdot, \cdot) is cosine similarity. Applying the softmax operator σ(⋅)\sigma(\cdot) with temperature parameter τ1\tau_1 yields two top-1 probability distributions: Pi=σ(S(xi)/τ1),Qi=σ(S(xi)′/τ1)P_i = \sigma(S(x_i) / \tau_1), \quad Q_i = \sigma(S(x_i)' / \tau_1)

    Ranking consistency is optimized by minimizing the Jensen-Shannon (JS) divergence between PiP_i and QiQ_i across all samples in the batch: Lconsistency=∑i=1NJS(Pi∥Qi)=12∑i=1N(Pilog⁡(2PiPi+Qi)+Qilog⁡(2QiPi+Qi))\mathcal{L}_{\text{consistency}} = \sum_{i=1}^N \text{JS}(P_i \parallel Q_i) = \frac{1}{2} \sum_{i=1}^N \left( P_i \log \left(\frac{2 P_i}{P_i + Q_i}\right) + Q_i \log \left(\frac{2 Q_i}{P_i + Q_i}\right) \right) JS divergence is utilized rather than Kullback-Leibler (KL) divergence because PiP_i and QiQ_i are symmetric representations of equal standing rather than ground-truth distributions.

  3. Knowl 3 — Listwise Ranking Distillation with ListNet and ListMLE

    model/method

    RankCSE distills listwise ranking knowledge from pre-trained teacher models into the student sentence encoder using teacher similarity scores as pseudo-ranking labels. Let S(xi)={ϕ(f(xi),f(xj)′)}j=1NS(x_i) = \{\phi(f(x_i), f(x_j)')\}_{j=1}^N denote the student similarity list and ST(xi)S^T(x_i) denote the pseudo-ranking similarity score list obtained from the teacher model for anchor sentence xix_i. Two listwise ranking distillation loss variants are supported:

    1. ListNet Distillation Loss (LListNet\mathcal{L}_{\text{ListNet}}): To make permutation probability computation tractable (O(N)O(N) instead of O(N!)O(N!)), ListNet aligns the top-1 probability distribution of the student with that of the teacher using cross-entropy: LListNet=−∑i=1Nσ(ST(xi)/τ3)⋅log⁡σ(S(xi)/τ2)\mathcal{L}_{\text{ListNet}} = -\sum_{i=1}^N \sigma(S^T(x_i) / \tau_3) \cdot \log \sigma(S(x_i) / \tau_2) where τ2\tau_2 and τ3\tau_3 are temperature hyperparameters and σ(⋅)\sigma(\cdot) is the softmax function. In practice, the score corresponding to the positive pair (xi,xi′)(x_i, x_i') is omitted when computing the top-1 distributions so that the loss focuses on ranking the negative samples.

    2. ListMLE Distillation Loss (LListMLE\mathcal{L}_{\text{ListMLE}}): ListMLE maximizes the likelihood of the ground-truth permutation πiT={πiT(j)}j=1N\pi_i^T = \{\pi_i^T(j)\}_{j=1}^N defined by the sorted indices of teacher similarity scores in descending order: LListMLE=−∑i=1Nlog⁡P(πiT∣S(xi),τ2)=−∑i=1N∑j=1Nlog⁡exp⁡(sπiT(j)/τ2)∑k=jNexp⁡(sπiT(k)/τ2)\mathcal{L}_{\text{ListMLE}} = -\sum_{i=1}^N \log P(\pi_i^T \mid S(x_i), \tau_2) = -\sum_{i=1}^N \sum_{j=1}^N \log \frac{\exp(s_{\pi_i^T(j)} / \tau_2)}{\sum_{k=j}^N \exp(s_{\pi_i^T(k)} / \tau_2)} where sms_m is the student's predicted similarity score for the mm-th sample.

  4. Knowl 4 — Multi-Teacher Ensembling for Pseudo-Ranking Distillation

    model/method

    To provide richer and more reliable pseudo-ranking labels for listwise distillation, RankCSE combines similarity scores from multiple pre-trained teacher models rather than relying on a single teacher. For an anchor sentence xix_i and two teacher models T1T_1 and T2T_2 (such as pre-trained base and large SimCSE checkpoints), the ensembled pseudo similarity score list ST(xi)S^T(x_i) is computed as the weighted average: ST(xi)=αS1T(xi)+(1−α)S2T(xi)S^T(x_i) = \alpha S_1^T(x_i) + (1 - \alpha) S_2^T(x_i) where S1T(xi)S_1^T(x_i) and S2T(xi)S_2^T(x_i) are the similarity score lists produced by teacher 1 and teacher 2 respectively, and α∈[0,1]\alpha \in [0, 1] is a balancing hyperparameter (set to α=1/3\alpha = 1/3 in default experiments). Aggregating predictions across multiple teachers reduces noise in pseudo-labels and exposes the student to complementary ranking signals.

  5. Knowl 5 — Empirical Performance on Semantic Textual Similarity (STS) Tasks

    empirical result

    Unsupervised sentence representation performance evaluated on 7 Semantic Textual Similarity (STS) benchmark test sets using Spearman's rank correlation (ρ×100\rho \times 100). All models are trained on 10610^6 randomly sampled unsupervised English Wikipedia sentences without labeled data, using cosine similarity between sentence embeddings for zero-shot evaluation.

    PLM Backbone / Method STS12 STS13 STS14 STS15 STS16 STS-B SICK-R Avg.
    BERT-base
    SimCSE 68.40 82.41 74.38 80.91 78.56 76.85 72.23 76.25
    DCLR 70.81 83.73 75.11 82.56 78.44 78.31 71.59 77.22
    ArcCSE 72.08 84.27 76.25 82.32 79.54 79.92 72.39 78.11
    DiffCSE 72.28 84.43 76.47 83.90 80.54 80.59 71.23 78.49
    PCL 72.84 83.81 76.52 83.06 79.32 80.01 73.38 78.42
    RankCSElistNet_{\text{listNet}} 74.38 85.97 77.51 84.46 81.31 81.46 75.26 80.05
    RankCSElistMLE_{\text{listMLE}} 75.66 86.27 77.81 84.74 81.10 81.80 75.13 80.36
    BERT-large
    SimCSE 70.88 84.16 76.43 84.50 79.76 79.26 73.88 78.41
    PCL 74.87 86.11 78.29 85.65 80.52 81.62 73.94 80.14
    RankCSElistNet_{\text{listNet}} 74.75 86.46 78.52 85.41 80.62 81.40 76.12 80.47
    RankCSElistMLE_{\text{listMLE}} 75.48 86.50 78.60 85.45 81.09 81.58 75.53 80.60
    RoBERTa-base
    SimCSE 70.16 81.77 73.24 81.36 80.65 80.22 68.56 76.57
    DiffCSE 70.05 83.43 75.49 82.81 82.12 82.38 71.19 78.21
    RankCSElistNet_{\text{listNet}} 72.91 85.72 76.94 84.52 82.59 83.46 71.94 79.73
    RankCSElistMLE_{\text{listMLE}} 73.20 85.95 77.17 84.82 82.58 83.08 71.88 79.81
    RoBERTa-large
    SimCSE 72.86 83.99 75.62 84.77 81.80 81.98 71.26 78.90
    PCL 74.08 84.36 76.42 85.49 81.76 82.79 71.51 79.49
    RankCSElistNet_{\text{listNet}} 73.47 85.77 78.07 85.65 82.51 84.12 73.73 80.47
    RankCSElistMLE_{\text{listMLE}} 73.20 85.83 78.00 85.63 82.67 84.19 73.64 80.45

    RankCSE consistently outperforms previous contrastive and post-processing methods across all backbones (p<0.005p < 0.005). RankCSE-BERT-base (80.36%80.36\%) surpasses SimCSE-BERT-large (78.41%78.41\%) by nearly 22 points.

  6. Knowl 6 — Empirical Performance on SentEval Transfer Tasks

    empirical result

    Performance of sentence representations on 7 standard SentEval downstream classification transfer tasks (MR, CR, SUBJ, MPQA, SST-2, TREC, MRPC) measured by classification accuracy (%) using a logistic regression classifier trained on top of frozen embeddings:

    PLM Backbone / Method MR CR SUBJ MPQA SST TREC MRPC Avg.
    BERT-base
    SimCSE 81.18 86.46 94.45 88.88 85.50 89.80 74.43 85.81
    DiffCSE 81.76 86.20 94.76 89.21 86.00 87.60 75.54 85.87
    RankCSElistNet_{\text{listNet}} 83.21 88.08 95.25 90.00 88.58 90.00 76.17 87.33
    RankCSElistMLE_{\text{listMLE}} 83.07 88.27 95.06 89.90 87.70 89.40 76.23 87.09
    BERT-large
    SimCSE 85.36 89.38 95.39 89.63 90.44 91.80 76.41 88.34
    RankCSElistNet_{\text{listNet}} 85.11 89.56 95.39 90.30 90.77 93.20 77.16 88.78
    RankCSElistMLE_{\text{listMLE}} 84.63 89.51 95.50 90.08 90.61 93.20 76.99 88.65
    RoBERTa-base
    SimCSE 81.04 87.74 93.28 86.94 86.60 84.60 73.68 84.84
    DiffCSE 82.42 88.34 93.51 87.28 87.70 86.60 76.35 86.03
    RankCSElistNet_{\text{listNet}} 83.53 89.22 94.07 88.97 89.95 89.20 76.52 87.35
    RankCSElistMLE_{\text{listMLE}} 83.32 88.61 94.03 88.88 89.07 90.80 76.46 87.31
    RoBERTa-large
    SimCSE 82.74 87.87 93.66 88.22 88.58 92.00 69.68 86.11
    PCL 84.47 89.06 94.60 89.26 89.02 94.20 74.96 87.94
    RankCSElistNet_{\text{listNet}} 84.47 89.51 94.65 89.87 89.46 93.00 75.88 88.12
    RankCSElistMLE_{\text{listMLE}} 84.61 89.27 94.47 89.99 89.73 92.60 74.43 87.87

    RankCSE outperforms all baseline models on transfer tasks across PLMs. RankCSElistNet_{\text{listNet}} achieves slightly higher transfer accuracy than RankCSElistMLE_{\text{listMLE}}, as top-1 probability distillation in ListNet is less sensitive to noisy full-permutation rankings than ListMLE.

  7. Knowl 7 — Ablation Study of Loss Components in RankCSE

    empirical result

    Ablation experiments on BERT-base analyzing the contribution of each loss function in the unified objective:

    Model Configuration STS (avg.) TR (avg.)
    SimCSE baseline 76.25 85.81
    RankCSElistNet_{\text{listNet}} 80.05 87.33
    w/o Lconsistency\mathcal{L}_{\text{consistency}} 79.56 86.80
    w/o LinfoNCE\mathcal{L}_{\text{infoNCE}} 79.72 86.91
    w/o Lconsistency,LinfoNCE\mathcal{L}_{\text{consistency}}, \mathcal{L}_{\text{infoNCE}} 79.41 86.76
    RankCSElistMLE_{\text{listMLE}} 80.36 87.09
    w/o Lconsistency\mathcal{L}_{\text{consistency}} 79.88 86.65
    w/o LinfoNCE\mathcal{L}_{\text{infoNCE}} 79.95 86.73
    w/o Lconsistency,LinfoNCE\mathcal{L}_{\text{consistency}}, \mathcal{L}_{\text{infoNCE}} 79.73 86.24
    RankCSE w/o Lrank\mathcal{L}_{\text{rank}} 76.93 85.97
    RankCSE w/o LinfoNCE,Lrank\mathcal{L}_{\text{infoNCE}}, \mathcal{L}_{\text{rank}} 73.74 85.56

    Key findings:

    1. Removing the ranking distillation loss Lrank\mathcal{L}_{\text{rank}} produces the largest drop in performance (from 80.05−80.3680.05-80.36 down to 76.9376.93 on STS), showing it is the core driver of fine-grained ranking quality.
    2. RankCSE with only Lrank\mathcal{L}_{\text{rank}} (79.41−79.7379.41-79.73) still outperforms the teacher baseline by distilling and ensembling knowledge from multiple teachers.
    3. Optimizing solely with Lconsistency\mathcal{L}_{\text{consistency}} drops STS score to 73.7473.74, because consistency alone cannot distinguish positive anchors from negative pairs.
    4. Combining all three loss terms yields optimal performance on both STS and transfer benchmarks.
  8. Knowl 8 — Impact of Teacher Architecture and Multi-Teacher Configurations on Ranking Distillation

    empirical result

    Average STS Spearman's correlation for RankCSE with BERT-base student under different teacher configurations:

    Teacher Model(s) RankCSEListNet_{\text{ListNet}} RankCSEListMLE_{\text{ListMLE}}
    SimCSEbase_{\text{base}} 77.48 77.75
    DiffCSEbase_{\text{base}} 78.87 79.06
    SimCSElarge_{\text{large}} 79.66 79.81
    SimCSEbase_{\text{base}} + DiffCSEbase_{\text{base}} 79.10 79.28
    SimCSEbase_{\text{base}} + SimCSElarge_{\text{large}} 80.05 80.36
    DiffCSEbase_{\text{base}} + SimCSElarge_{\text{large}} 80.20 80.47

    Key observations:

    1. Better teacher models produce better student representations (e.g., SimCSElarge_{\text{large}} teacher achieves 79.66/79.8179.66/79.81 vs. 77.48/77.7577.48/77.75 for SimCSEextbase_{ ext{base}}).
    2. Multi-teacher combinations consistently outperform their individual single-teacher components (e.g., SimCSEextbase_{ ext{base}} + SimCSEextlarge_{ ext{large}} reaches 80.05/80.3680.05/80.36, outperforming SimCSEextlarge_{ ext{large}} alone by +0.39+0.39 to +0.55+0.55).
    3. Distilling from heterogeneous teachers (DiffCSEextbase_{ ext{base}} + SimCSEextlarge_{ ext{large}}) achieves the highest STS performance (80.4780.47 with ListMLE).
  9. Knowl 9 — Evaluation on Fine-Grained Semantic Ranking Tasks (KCC and NDCG)

    empirical result

    To evaluate fine-grained semantic ranking fidelity beyond binary similarity classification, ranking tasks are constructed from each STS dataset: for each query sentence xix_i with more than three target sentence pairs {(xi,xij,yij)}j=1k\{(x_i, x_i^j, y_i^j)\}_{j=1}^k (k>3k > 3), target sentences are ranked by model similarity scores against ground-truth ratings yijy_i^j. Performance is evaluated with Kendall's correlation coefficient (KCC) and normalized discounted cumulative gain (NDCG) using a BERT-base encoder:

    Metric Method STS12 STS13 STS14 STS15 STS16 STS-B SICK-R Avg.
    KCC SimCSE 36.08 36.60 44.14 49.02 54.66 58.44 54.65 47.66
    DiffCSE 38.59 41.89 42.37 51.19 58.90 59.21 53.42 49.37
    RankCSE 42.79 46.26 44.53 52.00 57.21 63.64 57.40 51.98
    NDCG SimCSE 97.80 89.33 92.71 96.93 94.28 96.49 98.44 95.14
    DiffCSE 98.35 90.22 93.05 96.91 94.79 97.05 98.34 95.53
    RankCSE 98.20 92.27 93.46 97.21 95.24 97.45 98.67 96.07

    RankCSE achieves 51.9851.98 average KCC and 96.0796.07 average NDCG, significantly outperforming SimCSE (47.6647.66 KCC, 95.1495.14 NDCG) and DiffCSE (49.3749.37 KCC, 95.5395.53 NDCG), verifying that ranking consistency and ranking distillation capture continuous, fine-grained semantic distinctions.

  10. Knowl 10 — Alignment and Uniformity Trade-off in RankCSE Representation Space

    empirical result

    The representation space quality of sentence encoders is measured on the STS-B development set using alignment ℓalign\ell_{\text{align}} and uniformity ℓuniform\ell_{\text{uniform}} metrics: ℓalign≜E(x,x+)∼ppos∥f(x)−f(x+)∥2\ell_{\text{align}} \triangleq \mathbb{E}_{(x, x^+) \sim p_{\text{pos}}} \|f(x) - f(x^+)\|^2 ℓuniform≜log⁡Ex,y∼i.i.d.pdatae−2∥f(x)−f(y)∥2\ell_{\text{uniform}} \triangleq \log \mathbb{E}_{x, y \overset{i.i.d.}{\sim} p_{\text{data}}} e^{-2 \|f(x) - f(y)\|^2} where lower values indicate better representation geometry for both properties.

    Empirical comparison across BERT-base methods demonstrates:

    1. Untuned BERT embeddings suffer from severe anisotropy (high ℓalign≈0.70\ell_{\text{align}} \approx 0.70, high ℓuniform≈−1.3\ell_{\text{uniform}} \approx -1.3).
    2. SimCSE strongly optimizes uniformity (ℓuniform≈−2.5\ell_{\text{uniform}} \approx -2.5) but achieves higher alignment distance (ℓalign≈0.28\ell_{\text{align}} \approx 0.28).
    3. DiffCSE optimizes alignment (ℓalign≈0.18\ell_{\text{align}} \approx 0.18) at the expense of uniformity (ℓuniform≈−1.8\ell_{\text{uniform}} \approx -1.8).
    4. RankCSE achieves a superior balance (ℓalign≈0.12−0.19\ell_{\text{align}} \approx 0.12 - 0.19, ℓuniform≈−2.25−−2.40\ell_{\text{uniform}} \approx -2.25 - -2.40). By pulling semantically similar negative pairs closer through ranking distillation and consistency, RankCSE achieves significantly lower alignment error than SimCSE while maintaining a well-spread, uniform distribution on the hypersphere.
  11. Knowl 11 — Computational and Architectural Limitations of RankCSE

    limitation

    RankCSE presents two primary limitations:

    1. Computational Overhead during Training: Distilling listwise ranking knowledge requires generating pseudo-similarity rankings from pre-trained teacher models for each training batch. On a single NVIDIA Tesla A100 GPU (40GB), RankCSE-base requires ≈30\approx 30 minutes per epoch (120 minutes total for 4 epochs) compared to ≈20\approx 20 minutes for SimCSE (20 minutes total for 1 epoch), and RankCSE-large requires ≈55\approx 55 minutes per epoch (220 minutes total for 4 epochs) compared to 45 minutes for SimCSE.
    2. Heuristic Teacher Selection: In the multi-teacher formulation, teacher selection and weight balancing (e.g., combining SimCSEbase_{\text{base}} and SimCSElarge_{\text{large}} with α=1/3\alpha = 1/3) are determined heuristically rather than through a principled search method for optimal teacher model combinations.

Coverage note — All primary contributed methods, loss formulations, empirical benchmark evaluations across STS and transfer tasks, loss ablation studies, multi-teacher analyses, ranking metric evaluations, and stated limitations are included. Minor hyperparameter search grids and standard baseline descriptions are omitted.

References

  1. 1.Hervé Abdi. 2007. The kendall rank correlation coefficient. Encyclopedia of measurement and statistics, 2:508–510.
  2. 2.Eneko Agirre, Carmen Banea, Claire Cardie, Daniel M. Cer, Mona T. Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Iñigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, German Rigau, Larraitz Uria, and Janyce Wiebe. 2015. Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability. In Proceedings of the 9th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2015, Denver, Colorado, USA, June 4-5, 2015, pages 252–263. The Association for Computer Linguistics.
  3. 3.Eneko Agirre, Carmen Banea, Claire Cardie, Daniel M. Cer, Mona T. Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2014. Semeval-2014 task 10: Multilingual semantic textual similarity. In Proceedings of the 8th International Workshop on Semantic Evaluation, SemEval@COLING 2014, Dublin, Ireland, August 23-24, 2014, pages 81–91. The Association for Computer Linguistics.
  4. 4.Eneko Agirre, Carmen Banea, Daniel M. Cer, Mona T. Diab, Aitor Gonzalez-Agirre, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2016. Semeval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation. In Proceedings of the 10th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2016, San Diego, CA, USA, June 16-17, 2016, pages 497–511. The Association for Computer Linguistics.
  5. 5.Eneko Agirre, Daniel M. Cer, Mona T. Diab, and Aitor Gonzalez-Agirre. 2012. Semeval-2012 task 6: A pilot on semantic textual similarity. In Proceedings of the 6th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2012, Montréal, Canada, June 7-8, 2012, pages 385–393. The Association for Computer Linguistics.
  6. 6.Eneko Agirre, Daniel M. Cer, Mona T. Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013. *sem 2013 shared task: Semantic textual similarity. In Proceedings of the Second Joint Conference on Lexical and Computational Semantics, *SEM 2013, June 13-14, 2013, Atlanta, Georgia, USA, pages 32–43. Association for Computational Linguistics.
  7. 7.Christopher J. C. Burges, Robert Ragno, and Quoc Viet Le. 2006. Learning to rank with nonsmooth cost functions. In Advances in Neural Information Processing Systems 19, Proceedings of the Twentieth Annual Conference on Neural Information Processing Systems, Vancouver, British Columbia, Canada, December 4-7, 2006, pages 193–200. MIT Press.
  8. 8.Christopher J. C. Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Gregory N. Hullender. 2005. Learning to rank using gradient descent. In Machine Learning, Proceedings of the Twenty-Second International Conference (ICML 2005), Bonn, Germany, August 7-11, 2005, volume 119 of ACM International Conference Proceeding Series, pages 89–96. ACM.
  9. 9.Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. 2007. Learning to rank: from pairwise approach to listwise approach. In Machine Learning, Proceedings of the Twenty-Fourth International Conference (ICML 2007), Corvallis, Oregon, USA, June 20-24, 2007, volume 227 of ACM International Conference Proceeding Series, pages 129–136. ACM.
  10. 10.Fredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää Hellqvist, and Magnus Sahlgren. 2021. Semantic re-tuning with contrastive tension. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  11. 11.Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil. 2018. Universal sentence encoder for english. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018: System Demonstrations, Brussels, Belgium, October 31 - November 4, 2018, pages 169–174. Association for Computational Linguistics.
  12. 12.Daniel M. Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017. Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation, SemEval@ACL 2017, Vancouver, Canada, August 3-4, 2017, pages 1–14. Association for Computational Linguistics.
  13. 13.Ting Chen, Yizhou Sun, Yue Shi, and Liangjie Hong. 2017. On sampling strategies for neural network-based collaborative filtering. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Halifax, NS, Canada, August 13 - 17, 2017, pages 767–776. ACM.
  14. 14.Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljacic, Shang-Wen Li, Scott Yih, Yoon Kim, and James R. Glass. 2022. Diffcse: Difference-based contrastive learning for sentence embeddings. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 4207–4218. Association for Computational Linguistics.
  15. 15.Alexis Conneau and Douwe Kiela. 2018. Senteval: An evaluation toolkit for universal sentence representations. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation, LREC 2018, Miyazaki, Japan, May 7-12, 2018. European Language Resources Association (ELRA).
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics.
  17. 17.William B. Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing, IWP@IJCNLP 2005, Jeju Island, Korea, October 2005, 2005. Asian Federation of Natural Language Processing.
  18. 18.Kawin Ethayarajh. 2019. How contextual are contextualized word representations? comparing the geometry of bert, elmo, and GPT-2 embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 55–65. Association for Computational Linguistics.
  19. 19.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 6894–6910. Association for Computational Linguistics.
  20. 20.John M. Giorgi, Osvald Nitski, Bo Wang, and Gary D. Bader. 2021. Declutr: Deep contrastive learning for unsupervised textual representations. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 879–895. Association for Computational Linguistics.
  21. 21.Felix Hill, Kyunghyun Cho, and Anna Korhonen. 2016. Learning distributed representations of sentences from unlabelled data. In NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016, pages 1367–1377. The Association for Computational Linguistics.
  22. 22.Minqing Hu and Bing Liu. 2004. Mining and summarizing customer reviews. In Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Seattle, Washington, USA, August 22-25, 2004, pages 168–177. ACM.
  23. 23.Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Trans. Inf. Syst., 20(4):422–446.
  24. 24.Taeuk Kim, Kang Min Yoo, and Sang-goo Lee. 2021. Self-guided contrastive learning for BERT sentence representations. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 2528–2540. Association for Computational Linguistics.
  25. 25.Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Skip-thought vectors. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, pages 3294–3302.
  26. 26.Quoc V. Le and Tomás Mikolov. 2014. Distributed representations of sentences and documents. In Proceedings of the 31th International Conference on Machine Learning, ICML 2014, Beijing, China, 21-26 June 2014, volume 32 of JMLR Workshop and Conference Proceedings, pages 1188–1196. JMLR.org.
  27. 27.Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020. On the sentence embeddings from pre-trained language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 9119–9130. Association for Computational Linguistics.
  28. 28.Ping Li, Christopher J. C. Burges, and Qiang Wu. 2007. Mcrank: Learning to rank using multiple classification and gradient boosting. In Advances in Neural Information Processing Systems 20, Proceedings of the Twenty-First Annual Conference on Neural Information Processing Systems, Vancouver, British Columbia, Canada, December 3-6, 2007, pages 897–904. Curran Associates, Inc.
  29. 29.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  30. 30.Lajanugen Logeswaran and Honglak Lee. 2018. An efficient framework for learning sentence representations. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  31. 31.Shutian Ma, Chengzhi Zhang, and Daqing He. 2016. Document representation methods for clustering bilingual documents. In Creating Knowledge, Enhancing Lives through Information & Technology - Proceedings of the 2016 Annual Meeting of the Association for Information Science and Technology, ASIST 2016, Copenhagen, Denmark, October 14-18, 2016, volume 53 of Proc. Assoc. Inf. Sci. Technol., pages 1–10. Wiley.
  32. 32.Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014. A SICK cure for the evaluation of compositional distributional semantic models. In Proceedings of the Ninth International Conference on Language Resources and Evaluation, LREC 2014, Reykjavik, Iceland, May 26-31, 2014, pages 216–223. European Language Resources Association (ELRA).
  33. 33.Tomás Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, United States, pages 3111–3119.
  34. 34.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics, 21-26 July, 2004, Barcelona, Spain, pages 271–278. ACL.
  35. 35.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In ACL 2005, 43rd Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference, 25-30 June 2005, University of Michigan, USA, pages 115–124. The Association for Computer Linguistics.
  36. 36.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a Special Interest Group of the ACL, pages 1532–1543. ACL.
  37. 37.Przemyslaw Pobrotyn and Radoslaw Bialobrzeski. 2021. Neuralndcg: Direct optimisation of a ranking metric via differentiable relaxation of sorting. CoRR, abs/2102.07831.
  38. 38.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 3980–3990. Association for Computational Linguistics.
  39. 39.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013, 18-21 October 2013, Grand Hyatt Seattle, Seattle, Washington, USA, A meeting of SIGDAT, a Special Interest Group of the ACL, pages 1631–1642. ACL.
  40. 40.Jianlin Su, Jiarun Cao, Weijie Liu, and Yangyiwen Ou. 2021. Whitening sentence representations for better semantics and faster retrieval. CoRR, abs/2103.15316.
  41. 41.Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. CoRR, abs/1807.03748.
  42. 42.Maksims Volkovs and Richard S. Zemel. 2009. Boltzrank: learning to maximize expected ranking gain. In Proceedings of the 26th Annual International Conference on Machine Learning, ICML 2009, Montreal, Quebec, Canada, June 14-18, 2009, volume 382 of ACM International Conference Proceeding Series, pages 1089–1096. ACM.
  43. 43.Ellen M. Voorhees and Dawn M. Tice. 2000. Building a question answering test collection. In SIGIR 2000: Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, July 24-28, 2000, Athens, Greece, pages 200–207. ACM.
  44. 44.Kexin Wang, Nils Reimers, and Iryna Gurevych. 2021. TSDAE: using transformer-based sequential denoising auto-encoderfor unsupervised sentence embedding learning. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, pages 671–688. Association for Computational Linguistics.
  45. 45.Tongzhou Wang and Phillip Isola. 2020. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 9929–9939. PMLR.
  46. 46.Janyce Wiebe, Theresa Wilson, and Claire Cardie. 2005. Annotating expressions of opinions and emotions in language. Lang. Resour. Evaluation, 39(2-3):165–210.
  47. 47.Bohong Wu and Hai Zhao. 2022. Sentence representation learning with generative objective rather than contrastive objective. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 3356–3368. Association for Computational Linguistics.
  48. 48.Qiyu Wu, Chongyang Tao, Tao Shen, Can Xu, Xiubo Geng, and Daxin Jiang. 2022a. PCL: peer-contrastive learning with diverse augmentations for unsupervised sentence embeddings. CoRR, abs/2201.12093.
  49. 49.Xing Wu, Chaochen Gao, Zijia Lin, Jizhong Han, Zhongyuan Wang, and Songlin Hu. 2022b. Infocse: Information-aggregated contrastive learning of sentence embeddings. CoRR, abs/2210.06432.
  50. 50.Fen Xia, Tie-Yan Liu, Jue Wang, Wensheng Zhang, and Hang Li. 2008. Listwise approach to learning to rank: theory and algorithm. In Machine Learning, Proceedings of the Twenty-Fifth International Conference (ICML 2008), Helsinki, Finland, June 5-9, 2008, volume 307 of ACM International Conference Proceeding Series, pages 1192–1199. ACM.
  51. 51.Yuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang, Wei Wu, and Weiran Xu. 2021. Consert: A contrastive framework for self-supervised sentence representation transfer. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 5065–5075. Association for Computational Linguistics.
  52. 52.Yan Zhang, Ruidan He, Zuozhu Liu, Lidong Bing, and Haizhou Li. 2021. Bootstrapped unsupervised sentence representation learning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 5168–5180. Association for Computational Linguistics.
  53. 53.Yan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim, and Lidong Bing. 2020. An unsupervised sentence embedding method by mutual information maximization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 1601–1610. Association for Computational Linguistics.
  54. 54.Yuhao Zhang, Hongji Zhu, Yongliang Wang, Nan Xu, Xiaobo Li, and Binqiang Zhao. 2022. A contrastive framework for learning sentence representations from pairwise and triple-wise perspective in angular space. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 4892–4903. Association for Computational Linguistics.
  55. 55.Kun Zhou, Beichen Zhang, Xin Zhao, and Ji-Rong Wen. 2022. Debiased contrastive learning of unsupervised sentence representations. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 6120–6130. Association for Computational Linguistics.

Citation

MLA
Liu, J., et al. “RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 13785–802, https://doi.org/10.18653/v1/2023.acl-long.771.
APA
Liu, J., Liu, J., Wang, Q., Wang, J., Wu, W., Xian, Y., Zhao, D., Chen, K., & Yan, R. (2023). RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13785–13802. https://doi.org/10.18653/v1/2023.acl-long.771
Chicago
Liu, J., J. Liu, Q. Wang, et al. 2023. “RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13785–802. https://doi.org/10.18653/v1/2023.acl-long.771.
Harvard
Liu, J. et al. (2023) “RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 13785–13802. Available at: https://doi.org/10.18653/v1/2023.acl-long.771.
Vancouver
1. Liu J, Liu J, Wang Q, Wang J, Wu W, Xian Y, Zhao D, Chen K, Yan R (2023) RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 13785–13802

BibTeX

@inproceedings{liu-etal-2023-rankcse,
    title = "{R}ank{CSE}: Unsupervised Sentence Representations Learning via Learning to Rank",
    author = "Liu, Jiduan  and
      Liu, Jiahao  and
      Wang, Qifan  and
      Wang, Jingang  and
      Wu, Wei  and
      Xian, Yunsen  and
      Zhao, Dongyan  and
      Chen, Kai  and
      Yan, Rui",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.771/",
    doi = "10.18653/v1/2023.acl-long.771",
    pages = "13785--13802"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/