Built independently by an author, for readers. Read the story and support ChapterPal

keyword

attention networks

Attention networks are artificial neural network architectures that use attention mechanisms to dynamically focus on the most relevant parts of input data when performing a task. By computing learned numerical weights across different elements within an input, such as tokens in text, regions in images, or nodes in graphs, these models dynamically prioritize important features while filtering out less useful information. This capability enables the network to effectively capture long-range dependencies, model complex relational structures, and combine multimodal information, making it a foundational framework across diverse domains in machine learning and artificial intelligence.

2 items

Searching Large Neighborhoods for Integer Linear Programs with Contrastive Learning

Searching Large Neighborhoods for Integer Linear Programs with Contrastive Learning

Taoan Huang, Aaron M. Ferber, Yuandong Tian, Bistra Dilkina, Benoit Steiner

OrganizationsAnthropicMetaUniversity of Southern California

Why you should read this

Proposes a contrastive learning framework, CL-LNS, that trains graph attention networks on positive and negative neighborhood samples from local branching to learn fast, high-quality destroy heuristics for Large Neighborhood Search in integer linear programming.

Integer Linear Programs (ILPs) are powerful tools for modeling and solving a large number of combinatorial optimization problems. Recently, it has been shown that Large Neighborhood Search (LNS), as a heuristic algorithm, can find high-quality solutions to ILPs faster than Branch and Bound. However, how to find the right heuristics to maximize the performance of LNS remains an open problem. In this paper, we propose a novel approach, CL-LNS, that delivers state-of-the-art anytime performance on several ILP benchmarks measured by metrics including the primal gap, the primal integral, survival rates and the best performing rate. Specifically, CL-LNS collects positive and negative solution samples from an expert heuristic that is slow to compute and learns a more efficient one with contrastive learning. We use graph attention networks and a richer set of features to further improve its performance.

Added

2026-10-04

Towards VQA Models That Can Read

Towards VQA Models That Can Read

Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, Marcus Rohrbach

OrganizationsGeorgia Institute of TechnologyMeta

Why you should read this

Introduces the TextVQA dataset and the LoRRA model to enable visual question answering systems to read and reason about text embedded in everyday images.

Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing this problem. First, we introduce a new "TextVQA" dataset to facilitate progress on this important problem. Existing datasets either have a small proportion of questions about text (e.g., the VQA dataset) or are too small (e.g., the VizWiz dataset). TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Second, we introduce a novel model architecture that reads text in the image, reasons about it in the context of the image and the question, and predicts an answer which might be a deduction based on the text and the image or composed of the strings found in the image. Consequently, we call our approach Look, Read, Reason & Answer (LoRRA). We show that LoRRA outperforms existing state-of-the-art VQA models on our TextVQA dataset. We find that the gap between human performance and machine performance is significantly larger on TextVQA than on VQA 2.0, suggesting that TextVQA is well-suited to benchmark progress along directions complementary to VQA 2.0.

Added

2026-09-16