An Embodied Generalist Agent in 3D World

Jiangyong HuangSilong YongXiaojian MaXiongkun LinghuPuhao LiYan WangQing LiSong-Chun ZhuBaoxiong JiaSiyuan Huang

article2024ICML446 citations

Presents LEO, a multimodal generalist agent trained on unified vision-language-action sequences to perform complex 3D perception, spatial reasoning, robotic manipulation, and embodied navigation.

Listen

Modern artificial intelligence models have achieved remarkable success in general-purpose reasoning across two-dimensional images and text. However, they struggle to perceive, reason about, and physically act within the three-dimensional physical environments essential to real-world tasks and embodied intelligence. Developing agents capable of operating in the physical world has been severely hindered by the scarcity of large-scale 3D vision-language datasets, fragmented architectures, and the absence of unified learning strategies bridging perception and physical action.

The article aims to introduce and evaluate LEO, an embodied multi-modal generalist agent designed to perceive, ground, reason, plan, and act in 3D environments using a single unified model architecture. Specifically, it assesses whether an autoregressive language modeling framework can directly align 3D point clouds, egocentric 2D images, text instructions, and discrete embodied action tokens without relying on separate, task-specific output heads.

To accomplish this, the authors implemented a two-stage training scheme comprising 3D vision-language alignment followed by vision-language-action instruction tuning. They curated two large-scale datasets, LEO-align and LEO-instruct, leveraging automated data generation pipelines that prompt large models with structured 3D scene graphs and object-centric reasoning chains, followed by rigorous filtering. The model processes object-centric 3D point clouds and 2D images via dedicated encoders and adapters, feeding interleaved multi-modal tokens into a seven-billion-parameter language model tuned efficiently using low-rank adaptation. The system was systematically evaluated across diverse benchmarks covering 3D captioning, situated question answering, multi-round dialogue, task planning, simulated robotic manipulation, and object navigation.

The evaluation revealed several critical findings. First, LEO established new state-of-the-art results on 3D dense captioning and question answering benchmarks without requiring task-specific fine-tuning; on the Scan2Cap captioning benchmark, it achieved a CIDEr score of 72.4 compared to 66.9 from the best previous fine-tuned generalist model, while on ScanQA it scored 101.4 versus the prior best of 69.6. Second, the agent demonstrated robust competence in embodied acting, matching specialist performance in robotic manipulation and successfully zero-shot transferring navigation capabilities to novel environments. Third, scaling experiments proved that instruction-tuning loss decreases log-linearly with increased data scale and model parameters, conforming to established scaling laws. Finally, the analysis showed that while vision-language pretraining aids downstream embodied control, adding action tokens slightly degraded language performance, identifying an asymmetric transfer tradeoff.

These findings demonstrate that integrating 3D object-centric spatial representations directly into large foundation models is a viable path toward generalist physical agents. In practice, this unified architecture can lower development and deployment costs by replacing multiple brittle, task-specific modules with a single general-purpose system across robotics, assistive technologies, and ambient computing. However, decision-makers should note that joint training of action policies with language models requires careful data balancing to prevent degradation of conversational and reasoning capabilities.

Moving forward, practitioners should explore scaling up multi-modal 3D training datasets and adopt balanced data schemes with negative sampling to mitigate agent hallucinations. For embodied action, future engineering efforts should integrate policy recurrence into the architecture, as LEO's current feed-forward action policy struggles to match human demonstrations on complex exploratory paths. Further research must also prioritize safety validations and alignment protocols before deploying these embodied agents in unconstrained physical environments.

Confidence in these findings is well-supported by thorough quantitative ablations and comparative benchmarks across established indoor datasets. Nevertheless, limitations remain regarding generalization to complex, out-of-distribution real-world scenes and the simplified feed-forward control policy. Stakeholders should maintain cautious optimism and conduct targeted pilot studies in real physical setups before broad commercial implementation.

arXiv: 2311.12871
Cover for An Embodied Generalist Agent in 3D World

Abstract

Leveraging massive knowledge from large language models (LLMs), recent machine learning models show notable successes in general-purpose task solving in diverse domains such as computer vision and robotics. However, several significant challenges remain: (i) most of these models rely on 2D images yet exhibit a limited capacity for 3D input; (ii) these models rarely explore the tasks inherently defined in 3D world, e.g., 3D grounding, embodied reasoning and acting. We argue these limitations significantly hinder current models from performing real-world tasks and approaching general intelligence. To this end, we introduce LEO, an embodied multi-modal generalist agent that excels in perceiving, grounding, reasoning, planning, and acting in the 3D world. LEO is trained with a unified task interface, model architecture, and objective in two stages: (i) 3D vision-language (VL) alignment and (ii) 3D vision-language-action (VLA) instruction tuning. We collect large-scale datasets comprising diverse object-level and scene-level tasks, which require considerable understanding of and interaction with the 3D world. Moreover, we meticulously design an LLM-assisted pipeline to produce high-quality 3D VL data. Through extensive experiments, we demonstrate LEO’s remarkable proficiency across a wide spectrum of tasks, including 3D captioning, question answering, embodied reasoning, navigation and manipulation. Our ablative studies and scaling analyses further provide valuable insights for developing future embodied generalist agents. Code and data are available on project page.

Table of Contents

  • 1. Introduction
  • 2. Model
  • 2.1. Tokenization
  • 2.2. Token Embedding & LLM
  • 2.3. Training & Inference
  • 3. Datasets
  • 3.1. LEO-align: 3D Vision-Language Alignment
  • 3.2. LEO-instruct: Instruction Following in 3D world
  • 3.3. LLM-assisted 3D-language Data Generation
  • 4. Capabilities and Analyses
  • 4.1. 3D Vision-Language Understanding and Reasoning
  • 4.2. Scene-grounded Dialogue and Planning
  • 4.3. Embodied Action in 3D World
  • 4.4. More Insights into LEO
  • 4.5. Scaling Law Analysis
  • 5. Related Work
  • 6. Conclusions
  • Impact Statement
  • Acknowledgements
  • References
  • A. Qualitative Results
  • B. Data
  • B.1. More Details on LEO-align
  • B.2. More Details on LEO-instruct
  • B.3. Design of Seed Tasks for LLM-assisted 3D Data Generation
  • B.4. Prompts for LLM-assisted 3D Data Generation
  • B.5. Analysis of the Object-Centric Chain-of-Thought
  • B.6. Refinement Details
  • B.7. Subgraph Sampling
  • B.8. Scene-graph-based Prompting vs. Box-based Prompting
  • B.9. Additional Comparision Regarding Dataset Quality
  • B.10. Dataset Statistics
  • C. Data Examples
  • D. Model Details
  • D.1. Prompts
  • D.2. Feature Encoding
  • D.2.1. EMBODIMENT ENCODING
  • D.3. Action Tokenization
  • D.4. LLM Hyperparameters
  • E. Alignment Setup
  • F. Instruction-tuning Setup
  • G. Ablation Details
  • G.1. Object-centric Mask
  • G.2. Model Ablation
  • G.3. Dialogue and Planning Data
  • G.4. Data Balancing
  • H. Evaluation Details
  • H.1. 3D Question Answering
  • H.2. Embodied Navigation
  • I. Additional Results
  • I.1. Impact of Data Refinement
  • I.2. Data Comparison
  • I.3. Model Comparison
  • I.4. Embodied Acting
  • I.5. Scan2Cap
  • I.6. ScanQA
  • I.7. SQA3D

Knowls

  1. Knowl 1 — Unified Multi-Modal Sequence Architecture of the LEO Generalist Agent

    model/method

    LEO is an embodied multi-modal generalist agent designed to perform 3D scene perception, visual grounding, reasoning, task planning, and embodied acting within a single unified sequence-to-sequence model. The model unifies heterogeneous inputs and outputs into an interleaved sequence of tokens formatted as:

    You are...⏟system message    s2D(1),…,s2D(M)⏟2D image tokens (optional)    s3D(1),…,s3D(N)⏟object-centric 3D tokens    USER: … ASSISTANT:⏟instruction    sres(1),…,sres(T)⏟response\underbrace{\text{You are...}}_{\text{system message}} \;\; \underbrace{s^{(1)}_{\text{2D}}, \dots, s^{(M)}_{\text{2D}}}_{\text{2D image tokens (optional)}} \;\; \underbrace{s^{(1)}_{\text{3D}}, \dots, s^{(N)}_{\text{3D}}}_{\text{object-centric 3D tokens}} \;\; \underbrace{\text{USER: } \dots \text{ ASSISTANT:} }_{\text{instruction}} \;\; \underbrace{s^{(1)}_{\text{res}}, \dots, s^{(T)}_{\text{res}}}_{\text{response}}

    The architecture consists of:

    1. Text & Action Tokenizer: A SentencePiece tokenizer with a 32k32\text{k} subword vocabulary. Text tokens and mapped discrete action tokens are embedded via an embedding lookup table.
    2. Egocentric 2D Image Encoder: An OpenCLIP ConvNeXt-base vision backbone (pretrained on LAION-2B) that extracts MM 2D image token embeddings from an egocentric camera view.
    3. Object-Centric 3D Point Cloud Encoder: A PointNet++ encoder (pretrained on ScanNet object classification) that processes 10241024 sampled points per 3D object proposal (extracted via Mask3D) to produce initial point cloud embeddings for NN objects.
    4. Spatial Transformer: A 3-layer, 8-head modified Transformer that refines the initial 3D object point cloud embeddings by explicitly incorporating pairwise 3D spatial relations (distances and horizontal/vertical angles) into the self-attention weights.
    5. CLIP Semantic Guidance: The CLIP text encoder processes the instruction tokens to produce a global semantic feature vector, which is combined with every 2D image and 3D object token embedding via element-wise multiplication.
    6. Large Language Model Backbone: Vicuna-7B, augmented with Low-Rank Adaptation (LoRA) of rank r=16r = 16 and scaling factor α=16\alpha = 16 applied to all projection matrices (Wq,Wk,Wv,WoW_q, W_k, W_v, W_o in attention layers and Wgate,Wup,WdownW_{\text{gate}}, W_{\text{up}}, W_{\text{down}} in MLP layers). Out of approximately 7 billion total parameters, 142 million parameters (the 2D encoder, Spatial Transformer, MLP adapters, and LoRA parameters) are trainable, while the 3D PointNet++ encoder and the base LLM weights remain frozen.
  2. Knowl 2 — Two-Stage Generalist Training Scheme and Prefix Language Modeling Objective

    model/method

    LEO is trained using a two-stage learning framework formulated as prefix autoregressive language modeling. Given a batch B\mathcal{B} of token sequences ss, the objective optimizes the conditional likelihood of the response tokens given the prefix tokens:

    L(θ,B)=−∑b=1∣B∣∑t=1Tlog⁡pθ(sres(b,t)∣sres(b,<t),sprefix(b))\mathcal{L}(\theta, \mathcal{B}) = - \sum_{b=1}^{|\mathcal{B}|} \sum_{t=1}^T \log p_\theta\left(s_{\text{res}}^{(b,t)} \mid s_{\text{res}}^{(b,<t)}, s_{\text{prefix}}^{(b)}\right)

    where sprefix(b)s_{\text{prefix}}^{(b)} comprises the system message, optional 2D visual tokens, object-centric 3D tokens, and instruction tokens, while sres(b)s_{\text{res}}^{(b)} comprises the target response tokens.

    The training is structured into two stages:

    1. Stage 1: 3D Vision-Language (VL) Alignment (LEO-align): Aligns multi-modal 3D representations with natural language across three captioning tasks: (i) object-level captioning (single 3D objects from Objaverse/Cap3D), (ii) object-in-the-scene captioning (referring expressions of objects located within 3D scenes from ScanScribe, ReferIt3D, and 3RScan), and (iii) scene-level captioning (global scene descriptions derived from 3RScan scene graphs).
    2. Stage 2: 3D Vision-Language-Action (VLA) Instruction Tuning (LEO-instruct): Trains the agent to follow open-ended multi-modal instructions covering 3D dense captioning (Scan2Cap), 3D question answering (ScanQA, SQA3D, 3RScanQA), 3D dialogue (multi-turn interactions on 3RScan), scene-aware task planning (hierarchical action plans on 3RScan), embodied navigation (ObjectNav on Matterport3D/Habitat-web), and robotic manipulation (language-conditioned pick-and-place tasks from CLIPort).
  3. Knowl 3 — Spatial Transformer for Object-Centric 3D Spatial Relation Modeling

    model/method

    To model 3D geometric and spatial relationships among objects in a scene, LEO processes the point cloud embeddings of NN objects O∈RN×dO \in \mathbb{R}^{N \times d} using a Spatial Transformer that modulates standard multi-head self-attention with relative 3D coordinate geometry.

    For each pair of objects (Oi,Oj)(O_i, O_j) with 3D bounding box center coordinates ci,cj∈R3c_i, c_j \in \mathbb{R}^3, the Euclidean distance dij=∥ci−cj∥2d_{ij} = \|c_i - c_j\|_2 and the horizontal and vertical angles θh,θv\theta_h, \theta_v of the displacement vector connecting cic_i to cjc_j are computed to form a 5-dimensional relative spatial feature vector:

    fij=[dij,  sin⁡(θh),  cos⁡(θh),  sin⁡(θv),  cos⁡(θv)]f_{ij} = \left[d_{ij}, \; \sin(\theta_h), \; \cos(\theta_h), \; \sin(\theta_v), \; \cos(\theta_v)\right]

    A spatial projection vector gi=WSTOi∈R5g_i = W_S^T O_i \in \mathbb{R}^5 is computed from object embedding OiO_i with learnable matrix WSW_S. The scalar spatial attention bias ωijs\omega^s_{ij} is computed as the dot product:

    ωijs=giTfij\omega^s_{ij} = g_i^T f_{ij}

    Given the standard scaled dot-product attention score ωijo=(OiWQ)(OjWK)Tdh\omega^o_{ij} = \frac{(O_i W_Q)(O_j W_K)^T}{\sqrt{d_h}}, the modified spatial attention weight ωij\omega_{ij} between object ii and object jj is calculated as:

    ωij=σ(ωijs)exp⁡(ωijo)∑l=1Nσ(ωils)exp⁡(ωilo)\omega_{ij} = \frac{\sigma\left(\omega^s_{ij}\right) \exp\left(\omega^o_{ij}\right)}{\sum_{l=1}^N \sigma\left(\omega^s_{il}\right) \exp\left(\omega^o_{il}\right)}

    where σ(⋅)\sigma(\cdot) denotes the sigmoid function. The resulting re-weighted attention matrix is multiplied by the value representations V=OWVV = O W_V across 3 layers with 8 attention heads.

  4. Knowl 4 — Embodiment State Encoding and Discrete Action Space Tokenization

    model/method

    LEO incorporates an explicit embodiment representation and maps diverse robotic/embodied actions into discrete language subwords:

    1. Embodiment Token (ee): For tasks that require spatial awareness of an embodied agent (embodied navigation, situated reasoning on SQA3D, and object-in-the-scene captioning), a learnable self-object token ee is prepended to the 3D object sequence (e,s3D(1),…,s3D(N))(e, s^{(1)}_{\text{3D}}, \dots, s^{(N)}_{\text{3D}}). The embodiment token is assigned the agent's 3D position (from GPS or assumed object location) and orientation (from a compass sensor, encoded via Fourier features and mapped via a linear projection layer to feature vector rr). The full sequence of N+1N+1 tokens is processed jointly by the Spatial Transformer to model relative spatial relations between the agent and all scene objects.
    2. Discrete Action Tokenization: Continuous and discrete embodied actions are discretized and mapped into the least frequently used tokens in the SentencePiece vocabulary:
      • Embodied Navigation (Habitat ObjectNav): 4 reserved tokens represent the discrete primitive actions move forward, turn right, turn left, and stop.
      • Robotic Manipulation (CLIPort): Continuous 6-DoF end-effector pick-and-place poses are discretized into 516 discrete bins: 320 tokens for the x-axis position bins, 160 tokens for the y-axis position bins, and 36 tokens for the z-axis rotation bins.
      • Action execution is autoregressive: during navigation, LEO conditions on past 4 action tokens and current multi-modal observations to predict the next discrete action token.
  5. Knowl 5 — LLM-Assisted 3D-Language Data Generation Pipeline with O-CoT and Refinement

    model/method

    To overcome the scarcity of grounded 3D vision-language data, LEO introduces an automated data synthesis pipeline powered by ChatGPT that operates over 3D scene graphs from 3DSSG:

    1. Scene-Graph-Based Prompting: 3D scene graphs provide structured entity descriptions with class labels, bounding box coordinates, fine-grained object attributes (color, material, shape, state), and directional/topological spatial relations between adjacent objects.
    2. Subgraph Sampling: To ensure diverse descriptions and prevent the LLM from focusing solely on dominant objects in large scenes, subgraphs are sampled at dynamic preservation rates based on scene node counts (e.g., sampling rates 0.8–0.90.8\text{--}0.9 for 10–2010\text{--}20 nodes, down to 0.4–0.90.4\text{--}0.9 for >70>70 nodes).
    3. Object-Centric Chain-of-Thought (O-CoT): To prevent hallucinations during open-ended QA and dialogue generation, the LLM is instructed to explicitly generate intermediate reasoning thoughts containing the object name and instance ID (e.g., Thought: printer-8 or Thought: wardrobe-2, desk-7, bed-15) prior to emitting the final response.
    4. Rule-Based and Model-Based Refinement: Raw LLM responses undergo human-defined verification steps using regular expression matching:
      • Counting and Existence Filtering: Answers are automatically verified against ground-truth scene graph facts; incorrect counting or existence claims are fixed.
      • Negative Response Removal: Queries where the scene graph lacks information are discarded.
      • ID Stripping & Narration Rewriting: Intermediate thoughts and internal object IDs (e.g., dining table-33) are stripped, and sentences with residual formatting artifacts are rewritten by the LLM.

    On a test set of question types, raw LLM accuracy on Counting (57.4%), Existence (91.3%), and Non-existence (27.4%) improves to 78.0%, 93.4%, and 30.5% with O-CoT, and reaches 100.0% accuracy across all categories after refinement.

  6. Knowl 6 — Refined Exact Match Protocol for Open-Ended 3D Question Answering

    algorithm

    Standard strict Exact Match (Strict EM) misclassifies valid open-ended textual predictions generated by LLMs when predictions are semantically correct subsets or supersets of the reference string (e.g., predicting brown when the ground truth is dark brown, or 4 chairs when the ground truth is 4). LEO defines a Refined Exact Match (Refined EM) protocol that evaluates string containment after whitespace removal.

    Input: Model prediction string predpred, list of ground truth strings gtsgts
    Output: Boolean validation score (True if match, False otherwise)
    for each gtgt in gtsgts do
        if pred==gtpred == gt then
            return True
        end if
        pred_condensed = concatenate(split_whitespace(pred))
        gt_condensed = concatenate(split_whitespace(gt))
        if pred_condensed is substring of gt_condensed then
            return True
        end if
        if gt_condensed is substring of pred_condensed then
            return True
        end if
    end for
    return False
  7. Knowl 7 — Quantitative Performance of LEO on 3D Vision-Language and Embodied Reasoning Benchmarks

    data/table

    LEO was evaluated against single-task specialist models and task-specific fine-tuned models on Scan2Cap validation (using Mask3D object proposals with [email protected]), ScanQA validation, and SQA3D test splits. LEO achieves state-of-the-art performance across all three benchmarks within a single unified generalist model without task-specific heads or fine-tuning.

    Model Scan2Cap (val) ScanQA (val) SQA3D (test)
    C B-4 M R Sim C B-4 M R EM@1 EM@1
    Task-specific
    Scan2Cap 35.2 22.4 21.4 43.5 - - - - - - 41.0
    3DJCG 47.7 31.5 24.3 51.8 - - - - - - -
    Vote2Cap-DETR 61.8 34.5 26.2 54.4 - - - - - - -
    ScanRefer+MCAN - - - - - 55.4 7.9 11.5 30.0 18.6 -
    ClipBERT - - - - - - - - - - 43.3
    ScanQA - - - - - 64.9 10.1 13.1 33.3 21.1 47.2
    Task fine-tuned
    3D-VisTA 66.9 34.0 27.1 54.3 53.8 69.6 10.4 13.9 35.7 22.4 48.5
    3D-LLM (FlanT5) - - - - - 69.4 12.0 14.5 35.7 20.5 -
    LEO (Ours) 72.4 38.2 27.9 58.1 55.3 101.4 13.2 20.0 49.2 24.5 (47.6) 50.0 (52.4)

    Note: Metrics denote CIDEr (C), BLEU-4 (B-4), METEOR (M), ROUGE (R), Sentence Similarity (Sim), and Top-1 Exact Match (EM@1). Parentheses show results under the Refined Exact Match evaluation protocol.

  8. Knowl 8 — Performance of LEO on Robotic Manipulation and Embodied Navigation

    data/table

    LEO was evaluated on language-conditioned robotic manipulation (CLIPort benchmark) across seen and unseen object/color configurations, and on Object Goal Navigation (ObjNav) on Matterport3D (MP3D) and Habitat-Matterport 3D (HM3D) validation splits.

    Model separating-piles packing-google-objects-seq put-blocks-in-bowls
    seen unseen seen unseen seen unseen
    CLIP-only 90.2 71.0 95.8 57.8 97.7 44.5
    CLIPort (single) 98.0 75.2 96.2 71.9 100.0 25.0
    CLIPort (multi) 89.0 62.8 84.4 70.3 100.0 45.8
    LEO 98.8 75.2 76.6 79.8 86.2 35.2
    Model MP3D-val HM3D-val
    Success (%) ↑\uparrow SPL ↑\uparrow Success (%) ↑\uparrow SPL ↑\uparrow
    Habitat-web (shortest) 4.4 2.2 - -
    Habitat-web (demo) 35.4 10.2 - -
    ZSON (zero-shot) 15.3†15.3^\dagger 4.8†4.8^\dagger 25.5 12.6
    LEO 23.1 15.2 23.1†23.1^\dagger 19.1†\textbf{19.1}^\dagger

    Key findings:

    1. In robotic manipulation, LEO directly outputs discrete coordinate tokens without task-specific inductive spatial biases (such as 2D/3D heatmaps) and outperforms baselines on unseen out-of-distribution tasks (e.g., 79.8%79.8\% vs 71.9%71.9\% on unseen packing-google-objects-seq).
    2. In embodied navigation, LEO achieves a higher Success weighted by Path Length (SPL) of 15.215.2 on MP3D-val and 19.119.1 zero-shot on HM3D-val compared to specialist agents, demonstrating efficient path planning facilitated by global 3D object-centric spatial tokens.
  9. Knowl 9 — Ablations on 3D Alignment Stage, Generalist Scope, and VLA Co-Training Trade-offs

    data/table

    Ablation experiments evaluated LEO under different data configurations: omitting the Stage 1 alignment phase (w/o Align), training exclusively on ScanNet scenes (ScanNet), training without embodied acting tasks (w/o Act), and joint Vision-Language-Action training (VLA). Ground-truth object segments were used for these ablations.

    Configuration ScanNet 3RScan
    Scan2Cap (C) ScanQA (EM) SQA3D (EM) 3RQA (EM) 3RDialog (Sim) 3RPlan (Sim)
    w/o Align 62.8 22.7 (45.0) 50.9 (53.2) 49.7 (53.7) 73.0 80.3
    ScanNet Specialist 64.0 24.4 (49.2) 46.8 (49.5) 35.8 (50.0) 25.5 23.4
    w/o Act (VL default) 65.4 24.3 (48.5) 50.0 (52.5) 51.9 (57.4) 73.3 81.1
    VLA (Joint training) 65.3 25.0 (48.9) 46.2 (48.3) 51.3 (55.8) 72.3 77.2

    Conclusions:

    1. Pre-alignment: Omitting Stage 1 alignment (w/o Align) drops Scan2Cap CIDEr score from 65.465.4 to 62.862.8, highlighting that grounding spatial relationships and object attributes prior to instruction tuning is critical for captioning.
    2. Generalist vs. Specialist: Training strictly on ScanNet yields poor zero-shot transfer to novel 3RScan scenes (e.g., 3RPlan similarity drops from 81.181.1 to 23.423.4), proving the necessity of diverse generalist multi-task tuning.
    3. VLA Co-Training Trade-off: Co-training with embodied acting tasks (VLA) slightly degrades performance on 3D VL reasoning tasks (SQA3D drops from 50.050.0 to 46.246.2), attributed to the domain gap between natural language generation and motor action prediction and action dataset scale imbalances.
  10. Knowl 10 — Scaling Law Behavior and Object Existence Hallucination Mitigation in 3D LLMs

    empirical result

    LEO exhibits clear scaling properties with respect to model parameters, pre-alignment, and data balancing:

    1. Model and Data Scaling Law: Evaluation of LEO-instruct validation loss across training sample sizes (1.5×1041.5\times 10^4 to 12×10412\times 10^4) demonstrates that test loss decreases log-linearly with data scale. Comparing language model backbones shows:
      • Aligned Vicuna-7B achieves substantially lower test loss and superior benchmark performance over Aligned OPT-1.3B (e.g., ScanQA EM@1 increases from 20.320.3 to 24.324.3).
      • Aligned Vicuna-13B yields marginal additional gains over Vicuna-7B (ScanQA EM@1 of 23.423.4 vs 24.324.3), suggesting performance saturation at 7B parameters unless 3D training data scale is further expanded.
      • Pre-aligned models consistently outperform models trained from scratch without alignment across all data scales.
    2. Mitigating Object Existence Hallucinations: Pretrained 3D LLMs exhibit a positive confirmation bias, answering Yes to 98–100%98\text{--}100\% of object existence queries (Is there X in this room?), resulting in near-zero accuracy (1–16%1\text{--}16\%) on non-existent objects. By balancing the 3RScanQA tuning data with negative samples querying non-existent objects, accuracy on non-existent queries improves to 91%91\% on 3RScan and transfers zero-shot to ScanNet (81%81\% accuracy on non-existent objects, overall existence accuracy improving from 0.430.43 to 0.830.83).

Coverage note — None was omitted; all primary methodological contributions, training procedures, architectural components, datasets, empirical benchmarks, and ablation studies have been captured.

References

  1. 1.Achlioptas, P., Abdelreheem, A., Xia, F., Elhoseiny, M., and Guibas, L. Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes. In European Conference on Computer Vision (ECCV), 2020. 4, 7, 14
  2. 2.Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691, 2022. 8
  3. 3.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems (NeurIPS), 2022. 1, 3, 8
  4. 4.Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. Vqa: Visual question answering. In International Conference on Computer Vision (ICCV), 2015. 15
  5. 5.Azuma, D., Miyanishi, T., Kurita, S., and Kawanabe, M. Scanqa: 3d question answering for spatial scene understanding. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2, 6, 15
  6. 6.Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., et al. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023, 2023. 4
  7. 7.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021. 1
  8. 8.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022. 1, 8
  9. 9.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023. 1, 3, 7, 8, 30
  10. 10.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS), 2020. 1, 3
  11. 11.Cai, D., Zhao, L., Zhang, J., Sheng, L., and Xu, D. 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 6
  12. 12.Cai, S., Wang, Z., Ma, X., Liu, A., and Liang, Y. Openworld multi-task control through goal-aware representation learning and adaptive horizon prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13734–13744, 2023a. 8
  13. 13.Cai, S., Zhang, B., Wang, Z., Ma, X., Liu, A., and Liang, Y. Groot: Learning to follow instructions by watching gameplay videos. arXiv preprint arXiv:2310.08235, 2023b. 8
  14. 14.Chen, D. Z., Chang, A. X., and Nießner, M. Scanrefer: 3d object localization in rgb-d scans using natural language. In European Conference on Computer Vision (ECCV), 2020. 1
  15. 15.Chen, S., Guhur, P.-L., Tapaswi, M., Schmid, C., and Laptev, I. Language conditioned spatial relation reasoning for 3d object grounding. Advances in Neural Information Processing Systems (NeurIPS), 2022. 1, 3, 8, 25, 26
  16. 16.Chen, S., Zhu, H., Chen, X., Lei, Y., Yu, G., and Chen, T. End-to-end 3d dense captioning with vote2cap-detr. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 6
  17. 17.Chen, S., Chen, X., Zhang, C., Li, M., Yu, G., Fei, H., Zhu, H., Fan, J., and Chen, T. Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 31
  18. 18.Chen, Z., Gholami, A., Nießner, M., and Chang, A. X. Scan2cap: Context-aware dense captioning in rgb-d scans. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2, 4, 6, 15
  19. 19.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. Vicuna: An opensource chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/. 3, 8, 29
  20. 20.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022. 8
  21. 21.Dai, A., Chang, A. X., Savva, M., Halber, M., Funkhouser, T., and Nießner, M. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 1, 15, 25
  22. 22.Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. Instructblip: Towards general-purpose vision-language models with instruction tuning. arXiv preprint arXiv:2305.06500, 2023. 8
  23. 23.Deitke, M., Schwenk, D., Salvador, J., Weihs, L., Michel, O., VanderBilt, E., Schmidt, L., Ehsani, K., Kembhavi, A., and Farhadi, A. Objaverse: A universe of annotated 3d objects. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 14
  24. 24.Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al. Palm-e: An embodied multimodal language model. In International Conference on Machine Learning (ICML), 2023. 1, 8
  25. 25.Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A. Minedojo: Building open-ended embodied agents with internet-scale knowledge. Advances in Neural Information Processing Systems (NeurIPS), 2022. 8
  26. 26.Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., et al. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv preprint arXiv:2304.15010, 2023. 8
  27. 27.Gong, R., Huang, J., Zhao, Y., Geng, H., Gao, X., Wu, Q., Ai, W., Zhou, Z., Terzopoulos, D., Zhu, S.-C., et al. Arnold: A benchmark for language-grounded task learning with continuous states in realistic 3d scenes. In International Conference on Computer Vision (ICCV), 2023a. 8
  28. 28.Gong, R., Huang, Q., Ma, X., Vo, H., Durante, Z., Noda, Y., Zheng, Z., Zhu, S.-C., Terzopoulos, D., Fei-Fei, L., et al. Mindagent: Emergent gaming interaction. arXiv preprint arXiv:2309.09971, 2023b. 8
  29. 29.Gong, T., Lyu, C., Zhang, S., Wang, Y., Zheng, M., Zhao, Q., Liu, K., Zhang, W., Luo, P., and Chen, K. Multimodal-gpt: A vision and language model for dialogue with humans. arXiv preprint arXiv:2305.04790, 2023c. 8
  30. 30.Graepel, T., Minka, T., and Herbrich, R. T. A bayesian skill rating system. Advances in Neural Information Processing Systems, 19:569–576, 2007. 7, 29
  31. 31.Guo, J., Li, J., Li, D., Tiong, A. M. H., Li, B., Tao, D., and Hoi, S. C. From images to textual prompts: Zero-shot vqa with frozen large language models. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 8
  32. 32.Hong, Y., Zhen, H., Chen, P., Zheng, S., Du, Y., Chen, Z., and Gan, C. 3d-llm: Injecting the 3d world into large language models. arXiv preprint arXiv:2307.12981, 2023. 1, 4, 6, 8, 20, 31
  33. 33.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR), 2022. 3, 8, 27
  34. 34.Huang, J., Zhu, W. Y., Jia, B., Wang, Z., Ma, X., Li, Q., and Huang, S. Perceive, ground, reason, and act: A benchmark for general-purpose visual representation. arXiv preprint arXiv:2211.15402, 2022a. 1
  35. 35.Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., et al. Inner monologue: Embodied reasoning through planning with language models. In Conference on Robot Learning (CoRL), 2022b. 8
  36. 36.Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L. Vima: General robot manipulation with multimodal prompts. In International Conference on Machine Learning (ICML), 2023. 8
  37. 37.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. 2, 8
  38. 38.Kerr, J., Kim, C. M., Goldberg, K., Kanazawa, A., and Tancik, M. Lerf: Language embedded radiance fields. In International Conference on Computer Vision (ICCV), 2023. 8
  39. 39.Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023. 1, 8
  40. 40.Kudo, T. and Richardson, J. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226, 2018. 3
  41. 41.Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B. Human-level concept learning through probabilistic program induction. Science, 2015. 1
  42. 42.Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J. Building machines that learn and think like people. Behavioral and Brain Sciences, 2017. 1
  43. 43.Li, B., Zhang, Y., Chen, L., Wang, J., Pu, F., Yang, J., Li, C., and Liu, Z. Mimic-it: Multi-modal in-context instruction tuning. arXiv preprint arXiv:2306.05425, 2023a. 8
  44. 44.Li, B., Zhang, Y., Chen, L., Wang, J., Yang, J., and Liu, Z. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726, 2023b. 8
  45. 45.Li, C., Gan, Z., Yang, Z., Yang, J., Li, L., Wang, L., and Gao, J. Multimodal foundation models: From specialists to general-purpose assistants, 2023c. 1
  46. 46.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023d. 4, 8
  47. 47.Liu, F., Lin, K., Li, L., Wang, J., Yacoob, Y., and Wang, L. Aligning large multi-modal model with robust instruction tuning. arXiv preprint arXiv:2306.14565, 2023a. 8
  48. 48.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023b. 2, 3, 4, 8
  49. 49.Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. A convnet for the 2020s. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3, 26
  50. 50.Lu, J., Clark, C., Zellers, R., Mottaghi, R., and Kembhavi, A. Unified-io: A unified model for vision, language, and multi-modal tasks. In International Conference on Learning Representations (ICLR), 2023. 8
  51. 51.Luo, T., Rockwell, C., Lee, H., and Johnson, J. Scalable 3d captioning with pretrained models. arXiv preprint arXiv:2306.07279, 2023. 4, 8, 14
  52. 52.Ma, X., Yong, S., Zheng, Z., Li, Q., Liang, Y., Zhu, S.-C., and Huang, S. Sqa3d: Situated question answering in 3d scenes. In International Conference on Learning Representations (ICLR), 2023. 2, 4, 6, 15, 23
  53. 53.Majumdar, A., Aggarwal, G., Devnani, B., Hoffman, J., and Batra, D. Zson: Zero-shot object-goal navigation using multimodal goal embeddings. Advances in Neural Information Processing Systems (NeurIPS), 2022. 7
  54. 54.Mountcastle, V. B. An organizing principle for cerebral function: the unit module and the distributed system. The neurosciences. Fourth study program, 1979. 1
  55. 55.Mu, Y., Zhang, Q., Hu, M., Wang, W., Ding, M., Jin, J., Wang, B., Dai, J., Qiao, Y., and Luo, P. Embodiedgpt: Vision-language pre-training via embodied chain of thought. arXiv preprint arXiv:2305.15021, 2023. 8
  56. 56.OpenAI. Chatgpt. https://openai.com/blog/chatgpt/, 2022. 1, 2, 8, 16
  57. 57.OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 1, 2, 8, 20
  58. 58.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems (NeurIPS), 2022. 8
  59. 59.Peng, B., Li, C., He, P., Galley, M., and Gao, J. Instruction tuning with gpt-4. arXiv preprint arXiv:2304.03277, 2023a. 8
  60. 60.Peng, S., Genova, K., Jiang, C., Tagliasacchi, A., Pollefeys, M., Funkhouser, T., et al. Openscene: 3d scene understanding with open vocabularies. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023b. 8
  61. 61.Qi, C. R., Yi, L., Su, H., and Guibas, L. J. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in Neural Information Processing Systems (NeurIPS), 2017. 3, 25, 29
  62. 62.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021. 1, 27
  63. 63.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR), 2020. 3
  64. 64.Ramakrishnan, S. K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J., Undersander, E., Galuba, W., Westbury, A., Chang, A. X., et al. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai. arXiv preprint arXiv:2109.08238, 2021. 7, 30
  65. 65.Ramrakhya, R., Undersander, E., Batra, D., and Das, A. Habitat-web: Learning embodied object-search strategies from human demonstrations at scale. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2, 6, 7, 15, 27, 30, 32
  66. 66.Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al. A generalist agent. Transactions on Machine Learning Research (TMLR), 2022. 1, 2, 8
  67. 67.Reimers, N. and Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. In Annual Conference on Empirical Methods in Natural Language Processing (EMNLP), 2019. 6
  68. 68.Sanh, V., Webson, A., Raffel, C., Bach, S., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Raja, A., Dey, M., et al. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations (ICLR), 2022. 8
  69. 69.Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et al. Habitat: A platform for embodied ai research. In International Conference on Computer Vision (ICCV), 2019. 4, 15, 30
  70. 70.Schmidhuber, J. One big net for everything. arXiv preprint arXiv:1802.08864, 2018. 1
  71. 71.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems (NeurIPS), 2022. 1, 26
  72. 72.Schult, J., Engelmann, F., Hermans, A., Litany, O., Tang, S., and Leibe, B. Mask3d for 3d semantic instance segmentation. arXiv preprint arXiv:2210.03105, 2022. 3, 6, 28, 31
  73. 73.Shridhar, M., Manuelli, L., and Fox, D. Cliport: What and where pathways for robotic manipulation. In Conference on Robot Learning (CoRL), 2021. 2, 4, 6, 15, 27
  74. 74.Suglia, A., Gao, Q., Thomason, J., Thattai, G., and Sukhatme, G. Embodied bert: A transformer model for embodied, language-guided visual task completion. arXiv preprint arXiv:2108.04927, 2021. 30
  75. 75.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023. 8
  76. 76.Tsimpoukelli, M., Menick, J. L., Cabi, S., Eslami, S., Vinyals, O., and Hill, F. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems (NeurIPS), 2021. 1, 8
  77. 77.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in Neural Information Processing Systems (NeurIPS), 2017. 1, 25
  78. 78.Wald, J., Avetisyan, A., Navab, N., Tombari, F., and Nießner, M. Rio: 3d object instance re-localization in changing indoor environments. In International Conference on Computer Vision (ICCV), 2019. 1, 15
  79. 79.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023a. 8
  80. 80.Wang, X., Wang, W., Cao, Y., Shen, C., and Huang, T. Images speak in images: A generalist painter for in-context visual learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023b. 8
  81. 81.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language model with self generated instructions. In Annual Meeting of the Association for Computational Linguistics (ACL), 2023c. 8
  82. 82.Wang, Z., Cai, S., Liu, A., Ma, X., and Liang, Y. Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents. arXiv preprint arXiv:2302.01560, 2023d. 8
  83. 83.Wang, Z., Huang, H., Zhao, Y., Zhang, Z., and Zhao, Z. Chat-3d: Data-efficiently tuning large language model for universal dialogue of 3d scenes. arXiv preprint arXiv:2308.08769, 2023e. 4, 8
  84. 84.Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. In International Conference on Learning Representations (ICLR), 2022. 8
  85. 85.Wu, S.-C., Wald, J., Tateno, K., Navab, N., and Tombari, F. Scenegraphfusion: Incremental 3d scene graph prediction from rgb-d sequences. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 4, 17
  86. 86.Xu, R., Wang, X., Wang, T., Chen, Y., Pang, J., and Lin, D. Pointllm: Empowering large language models to understand point clouds. arXiv preprint arXiv:2308.16911, 2023. 8, 29
  87. 87.Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178, 2023. 8
  88. 88.Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Sheng, L., Bai, L., Huang, X., Wang, Z., et al. Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark. arXiv preprint arXiv:2306.06687, 2023. 4, 8
  89. 89.Yu, X., Tang, L., Rao, Y., Huang, T., Zhou, J., and Lu, J. Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 29
  90. 90.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022. 8, 29
  91. 91.Zhao, H., Cai, Z., Si, S., Ma, X., An, K., Chen, L., Liu, Z., Wang, S., Han, W., and Chang, B. Mmicl: Empowering vision-language model with multi-modal in-context learning. arXiv preprint arXiv:2309.07915, 2023. 8
  92. 92.Zhao, L., Cai, D., Sheng, L., and Xu, D. 3dvg-transformer: Relation modeling for visual grounding on point clouds. In International Conference on Computer Vision (ICCV), 2021. 1, 8
  93. 93.Zhu, D., Chen, J., Haydarov, K., Shen, X., Zhang, W., and Elhoseiny, M. Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions. arXiv preprint arXiv:2303.06594, 2023a. 8
  94. 94.Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023b. 8
  95. 95.Zhu, Y., Gao, T., Fan, L., Huang, S., Edmonds, M., Liu, H., Gao, F., Zhang, C., Qi, S., Wu, Y. N., et al. Dark, beyond deep: A paradigm shift to cognitive ai with humanlike common sense. Engineering, 2020. 1
  96. 96.Zhu, Z., Ma, X., Chen, Y., Deng, Z., Huang, S., and Li, Q. 3d-vista: Pre-trained transformer for 3d vision and text alignment. In International Conference on Computer Vision (ICCV), 2023c. 1, 3, 4, 6, 8, 14, 22

Citation

MLA
Huang, J., et al. “An Embodied Generalist Agent in 3D World”. arXiv, 2023, http://arxiv.org/abs/2311.12871v3.
APA
Huang, J., Yong, S., Ma, X., Linghu, X., Li, P., Wang, Y., Li, Q., Zhu, S.-C., Jia, B., & Huang, S. (2023). An Embodied Generalist Agent in 3D World. arXiv. http://arxiv.org/abs/2311.12871v3
Chicago
Huang, J., S. Yong, X. Ma, et al. 2023. “An Embodied Generalist Agent in 3D World”. arXiv. http://arxiv.org/abs/2311.12871v3.
Harvard
Huang, J. et al. (2023) “An Embodied Generalist Agent in 3D World”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.12871v3.
Vancouver
1. Huang J, Yong S, Ma X, Linghu X, Li P, Wang Y, Li Q, Zhu S-C, Jia B, Huang S (2023) An Embodied Generalist Agent in 3D World. arXiv

BibTeX

@article{huang2023embodied,
  title = {An Embodied Generalist Agent in 3D World},
  author = {Huang, Jiangyong and Yong, Silong and Ma, Xiaojian and Linghu, Xiongkun and Li, Puhao and Wang, Yan and Li, Qing and Zhu, Song-Chun and Jia, Baoxiong and Huang, Siyuan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.12871v3},
  eprint = {2311.12871}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/