Built independently by an author, for readers. Read the story and support ChapterPal

keyword

natural language

Natural language is any mode of communication, such as spoken or written language, that has developed organically within human communities through use and cultural evolution rather than being intentionally or artificially constructed. Unlike formal languages, which include computer programming code and mathematical notations governed by strict, unambiguous syntax, natural languages are characterized by flexible grammatical rules, diverse vocabularies, idioms, and contextual nuance. In computational contexts and artificial intelligence, natural language represents the unstructured textual or verbal information used by humans to convey meaning, pose questions, and describe concepts, forming the basis for tasks such as automated text interpretation, dialogue generation, and semantic reasoning.

3 items

Text2Loc: 3D Point Cloud Localization from Natural Language

Text2Loc: 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Zifeng Ding, João F. Henriques, Daniel Cremers

OrganizationsLudwig Maximilian University of MunichMunich Center for Machine LearningTechnical University of MunichUniversity of Oxford

Why you should read this

Proposes Text2Loc, a coarse-to-fine framework that localizes natural language descriptions within city-scale 3D point clouds by combining a hierarchical transformer for cross-sentence context with a matching-free fine localization network.

We tackle the problem of 3D point cloud localization based on a few natural linguistic descriptions and introduce a novel neural network, Text2Loc, that fully interprets the semantic relationship between points and text. Text2Loc follows a coarse-to-fine localization pipeline: text-submap global place recognition, followed by fine localization. In global place recognition, relational dynamics among each textual hint are captured in a hierarchical transformer with max-pooling (HTM), whereas a balance between positive and negative pairs is maintained using text-submap contrastive learning. Moreover, we propose a novel matching-free fine localization method to further refine the location predictions, which completely removes the need for complicated text-instance matching and is lighter, faster, and more accurate than previous methods. Extensive experiments show that Text2Loc improves the localization accuracy by up to 2× over the state-of-the-art on the KITTI360Pose dataset. Our project page is publicly available at https://yan-xia.github.io/projects/text2loc/.

Added

2026-09-26

Program Synthesis with Large Language Models

Program Synthesis with Large Language Models

Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, Charles Sutton

OrganizationsGoogleMassachusetts Institute of Technology

Why you should read this

Introduces the MBPP benchmark for Python code synthesis, demonstrating that large language model performance scales log-linearly with model size and that conversational human feedback cuts error rates in half.

This paper explores the limits of the current generation of large language models for program synthesis in general purpose programming languages. We evaluate a collection of such models (with between 244M and 137B parameters) on two new benchmarks, MBPP and MathQA-Python, in both the few-shot and fine-tuning regimes. Our benchmarks are designed to measure the ability of these models to synthesize short Python programs from natural language descriptions. The Mostly Basic Programming Problems (MBPP) dataset contains 974 programming tasks, designed to be solvable by entry-level programmers. The MathQA-Python dataset, a Python version of the MathQA benchmark, contains 23914 problems that evaluate the ability of the models to synthesize code from more complex text. On both datasets, we find that synthesis performance scales log-linearly with model size. Our largest models, even without finetuning on a code dataset, can synthesize solutions to 59.6 percent of the problems from MBPP using few-shot learning with a well-designed prompt. Fine-tuning on a held-out portion of the dataset improves performance by about 10 percentage points across most model sizes. On the MathQA-Python dataset, the largest fine-tuned model achieves 83.8 percent accuracy. Going further, we study the model's ability to engage in dialog about code, incorporating human feedback to improve its solutions. We find that natural language feedback from a human halves the error rate compared to the model's initial prediction. Additionally, we conduct an error analysis to shed light on where these models fall short and what types of programs are most difficult to generate. Finally, we explore the semantic grounding of these models by fine-tuning them to predict the results of program execution. We find that even our best models are generally unable to predict the output of a program given a specific input.

Added

2026-09-16

PIQA: Reasoning about Physical Commonsense in Natural Language

PIQA: Reasoning about Physical Commonsense in Natural Language

Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, Yejin Choi

OrganizationsAllen Institute for AICarnegie Mellon UniversityMicrosoftUniversity of Washington

Why you should read this

Introduces the PIQA benchmark to evaluate physical commonsense reasoning in natural language processing, exposing critical knowledge gaps where state-of-the-art models fail to match human understanding of everyday physical interactions and object affordances.

To apply eyeshadow without a brush, should I use a cotton swab or a toothpick? Questions requiring this kind of physical commonsense pose a challenge to today's natural language understanding systems. While recent pretrained models (such as BERT) have made progress on question answering over more abstract domains - such as news articles and encyclopedia entries, where text is plentiful - in more physical domains, text is inherently limited due to reporting bias. Can AI systems learn to reliably answer physical common-sense questions without experiencing the physical world? In this paper, we introduce the task of physical commonsense reasoning and a corresponding benchmark dataset Physical Interaction: Question Answering or PIQA. Though humans find the dataset easy (95% accuracy), large pretrained models struggle (77%). We provide analysis about the dimensions of knowledge that existing models lack, which offers significant opportunities for future research.

Added

2026-09-11