Built independently by an author, for readers. Read the story and support ChapterPal

topic

surface acoustic waves (SAW, surface acoustic wave)

A surface acoustic wave is an acoustic wave that propagates along the surface of an elastic solid, with its displacement amplitude confined primarily to the material boundary and decaying exponentially with depth into the substrate. In electronics, signal processing, and computing hardware, surface acoustic waves are typically generated and detected on piezoelectric materials using interdigital transducers, which convert alternating electrical signals into mechanical surface vibrations and vice versa. Because acoustic waves travel orders of magnitude slower than electromagnetic waves, surface acoustic waves enable complex high-frequency signal manipulation, filtering, and sensing to occur within compact physical components, such as radio-frequency bandpass filters, resonators, delay lines, and tactile sensors.

19 items

Smoothed Adaptive Weighting for Imbalanced Semi-Supervised Learning: Improve Reliability Against Unknown Distribution Data

Smoothed Adaptive Weighting for Imbalanced Semi-Supervised Learning: Improve Reliability Against Unknown Distribution Data

Zhengfeng Lai, Chao Wang, Henrry Gunawan, Sen-Ching S. Cheung, Chen-Nee Chuah

OrganizationsSouthern University of Science and TechnologyUniversity of California, DavisUniversity of Kentucky

Why you should read this

Proposes a smoothed adaptive weighting framework that dynamically adjusts consistency loss based on per-class learning difficulty, enabling semi-supervised models to handle severely imbalanced data without prior knowledge of the unlabeled distribution.

Despite recent promising results on semi-supervised learning (SSL), data imbalance, particularly in the unlabeled dataset, could significantly impact the training performance of a SSL algorithm if there is a mismatch between the expected and actual class distributions. The efforts on how to construct a robust SSL framework that can effectively learn from datasets with unknown distributions remain limited. We first investigate the feasibility of adding weights to the consistency loss and then we verify the necessity of smoothed weighting schemes. Based on this study, we propose a self-adaptive algorithm, named Smoothed Adaptive Weighting (SAW). SAW is designed to enhance the robustness of SSL by estimating the learning difficulty of each class and synthesizing the weights in the consistency loss based on such estimation. We show that SAW can complement recent consistency-based SSL algorithms and improve their reliability on various datasets including three standard datasets and one gigapixel medical imaging application without making any assumptions about the distribution of the unlabeled set.

Added

2026-10-03

Mimetic Initialization of Self-Attention Layers

Mimetic Initialization of Self-Attention Layers

Asher Trockman, J. Zico Kolter

OrganizationsBosch Center for AICarnegie Mellon University

Why you should read this

Proposes a simple, learning-free weight initialization strategy for self-attention layers that mimics weight patterns observed in pretrained models, boosting classification accuracy by up to 5% when training vanilla Vision Transformers from scratch on small and medium datasets.

It is notoriously difficult to train Transformers on small datasets; typically, large pre-trained models are instead used as the starting point. We explore the weights of such pre-trained Transformers (particularly for vision) to attempt to find reasons for this discrepancy. Surprisingly, we find that simply initializing the weights of self-attention layers so that they “look” more like their pre-trained counterparts allows us to train vanilla Transformers faster and to higher final accuracies, particularly on vision tasks such as CIFAR-10 and ImageNet classification, where we see gains in accuracy of over 5% and 4%, respectively. Our initialization scheme is closed form, learning-free, and very simple: we set the product of the query and key weights to be approximately the identity, and the product of the value and projection weights to approximately the negative identity. As this mimics the patterns we saw in pre-trained Transformers, we call the technique mimetic initialization.

Added

2026-09-26

Introduction to Sociology - 2nd Canadian Edition

Introduction to Sociology - 2nd Canadian Edition

William Little

Introduction to Sociology adheres to the scope and sequence of a typical introductory sociology course. In addition to comprehensive coverage of core concepts, foundational scholars, and emerging theories, we have incorporated section reviews with engaging questions, discussions that help students apply the sociological imagination, and features that draw learners into the discipline in meaningful ways. Although this text can be modified and reorganized to suit your needs, the standard version is organized so that topics are introduced conceptually, with relevant, everyday experiences. For the student, this book is based on the teaching and research experience of numerous sociologists. In today’s global socially networked world, the topic of Sociology is more relevant than ever before. We hope that through this book, students will learn how simple, everyday human actions and interactions can change the world. In this book, you will find applications of Sociology concepts that are relevant, current, and balanced. For instructors, this text is intended for a one-semester introductory course and includes these features: Sociological Research: Highlights specific current and relevant research studies. Sociology in the Real World: Ties chapter content to student life and discusses sociology in terms of the everyday. Big Picture: Features present sociological concepts at a national or international level. Case Study: Describes real-life people whose experiences relate to chapter content. Social Policy and Debate: Discusses political issues that relate to chapter content. Section Summaries distill the information in each section for both students and instructors down to key, concise points addressed in the section. Key Terms are bold and are followed by a definition in context. Definitions of key terms are also listed in the Key Terms, which appears at the end of each chapter. Section Quizzes provide opportunities to apply and test the information students learn throughout each section. Both multiple-choice and short-response questions feature a variety of question types and range of difficulty. Further Research : This feature helps students further explore the section topic and offers related research topics that could be explored.

Added

2026-09-25

Creative Commons License
The Principles of Deep Learning Theory

The Principles of Deep Learning Theory

Daniel A. Roberts, Sho Yaida, Boris Hanin

OrganizationsMassachusetts Institute of TechnologyMeta

Why you should read this

Establishes an effective theory framework for finite-width neural networks that explains how the depth-to-width ratio governs representation learning, gradient propagation, and optimal architecture design from first principles.

This book develops an effective theory approach to understanding deep neural networks of practical relevance. Beginning from a first-principles component-level picture of networks, we explain how to determine an accurate description of the output of trained networks by solving layer-to-layer iteration equations and nonlinear learning dynamics. A main result is that the predictions of networks are described by nearly-Gaussian distributions, with the depth-to-width aspect ratio of the network controlling the deviations from the infinite-width Gaussian description. We explain how these effectively-deep networks learn nontrivial representations from training and more broadly analyze the mechanism of representation learning for nonlinear models. From a nearly-kernel-methods perspective, we find that the dependence of such models' predictions on the underlying learning algorithm can be expressed in a simple and universal way. To obtain these results, we develop the notion of representation group flow (RG flow) to characterize the propagation of signals through the network. By tuning networks to criticality, we give a practical solution to the exploding and vanishing gradient problem. We further explain how RG flow leads to near-universal behavior and lets us categorize networks built from different activation functions into universality classes. Altogether, we show that the depth-to-width ratio governs the effective model complexity of the ensemble of trained networks. By using information-theoretic techniques, we estimate the optimal aspect ratio at which we expect the network to be practically most useful and show how residual connections can be used to push this scale to arbitrary depths. With these tools, we can learn in detail about the inductive bias of architectures, hyperparameters, and optimizers.

Added

2026-09-15

License

Published with permission

Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Ofir Press, Noah A. Smith, Mike Lewis

OrganizationsAllen Institute for AIMetaUniversity of Washington

Why you should read this

Introduces a static positional bias that allows Transformers to generalize to sequence lengths far beyond those encountered during training.

Since the introduction of the transformer model by Vaswani et al. (2017), a fundamental question has yet to be answered: how does a model achieve extrapolation at inference time for sequences that are longer than it saw during training? We first show that extrapolation can be enabled by simply changing the position representation method, though we find that current methods do not allow for efficient extrapolation. We therefore introduce a simpler and more efficient position method, Attention with Linear Biases (ALiBi). ALiBi does not add positional embeddings to word embeddings; instead, it biases query-key attention scores with a penalty that is proportional to their distance. We show that this method trains a 1.3 billion parameter model on input sequences of length 1024 that extrapolates to input sequences of length 2048, achieving the same perplexity as a sinusoidal position embedding model trained on inputs of length 2048 but training 11% faster and using 11% less memory. ALiBi's inductive bias towards recency also leads it to outperform multiple strong position methods on the WikiText-103 benchmark.

Added

2026-02-09