Built independently by an author, for readers. Read the story and support ChapterPal

keyword

shape representation

Shape representation refers to the mathematical or computational method used to encode, describe, and store the spatial geometry and boundary structure of an object in two or three dimensions. In fields such as computer vision, computer graphics, and digital image processing, it provides a structured framework that allows algorithms to process, analyze, reconstruct, and generate physical forms. Common approaches range from explicit models, such as polygon meshes, point clouds, and parametric curves, to implicit formulations, including continuous signed distance fields and volumetric occupancy grids, as well as learned neural feature embeddings and multi-view descriptors. The choice of shape representation determines how accurately local surface details, topological variations, and global spatial relationships are preserved during computational tasks such as object recognition, segmentation, 3D reconstruction, and rendering.

8 items

Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing

Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing

Thiemo Alldieck, Mihai Zanfir, Cristian Sminchisescu

OrganizationsGoogle

Why you should read this

Presents PHORHUM, an end-to-end deep learning framework that reconstructs photorealistic 3D clothed humans from a single RGB image by simultaneously estimating detailed geometry, scene illumination, and unshaded surface albedo without requiring foreground masks.

We present PHORHUM, a novel, end-to-end trainable, deep neural network methodology for photorealistic 3D human reconstruction given just a monocular RGB image. Our pixel-aligned method estimates detailed 3D geometry and, for the first time, the unshaded surface color together with the scene illumination. Observing that 3D supervision alone is not sufficient for high fidelity color reconstruction, we introduce patch-based rendering losses that enable reliable color reconstruction on visible parts of the human, and detailed and plausible color estimation for the non-visible parts. Moreover, our method specifically addresses methodological and practical limitations of prior work in terms of representing geometry, albedo, and illumination effects, in an end-to-end model where factors can be effectively disentangled. In extensive experiments, we demonstrate the versatility and robustness of our approach. Our state-of-the-art results validate the method qualitatively and for different metrics, for both geometric and color reconstruction.

Added

2026-10-06

Surface Representation for Point Clouds

Surface Representation for Point Clouds

Haoxi Ran, Jun Liu, Chengjie Wang

Why you should read this

Proposes RepSurf, a lightweight, plug-and-play surface representation that explicitly models local geometry using graphics-inspired triangle and umbrella primitives to boost accuracy across 3D classification, segmentation, and detection tasks with negligible computational overhead.

Most prior work represents the shapes of point clouds by coordinates. However, it is insufficient to describe the local geometry directly. In this paper, we present RepSurf (representative surfaces), a novel representation of point clouds to explicitly depict the very local structure. We explore two variants of RepSurf, Triangular RepSurf and Umbrella RepSurf inspired by triangle meshes and umbrella curvature in computer graphics. We compute the representations of RepSurf by predefined geometric priors after surface reconstruction. RepSurf can be a plug-and-play module for most point cloud models thanks to its free collaboration with irregular points. Based on a simple baseline of PointNet++ (SSG version), Umbrella RepSurf surpasses the previous state-of-the-art by a large margin for classification, segmentation and detection on various benchmarks in terms of performance and efficiency. With an increase of around 0.008M number of parameters, 0.04G FLOPs, and 1.12ms inference time, our method achieves 94.7% (+0.5%) on ModelNet40, and 84.6% (+1.8%) on ScanObjectNN for classification, while 74.3% (+0.8%) mIoU on S3DIS 6-fold, and 70.0% (+1.6%) mIoU on ScanNet for segmentation. For detection, previous state-of-the-art detector with our RepSurf obtains 71.2% (+2.1%) mAP25, 54.8% (+2.0%) mAP50 on ScanNetV2, and 64.9% (+1.9%) mAP25, 47.7% (+2.5%) mAP50 on SUN RGB-D. Our lightweight Triangular RepSurf performs its excellence on these benchmarks as well. The code is publicly available at https://github.com/hancyran/RepSurf.

Added

2026-10-05

RBGNet: Ray-based Grouping for 3D Object Detection

RBGNet: Ray-based Grouping for 3D Object Detection

Haiyang Wang, Shaoshuai Shi, Ze Yang, Rongyao Fang, Qi Qian, Hongsheng Li, Bernt Schiele, Liwei Wang

OrganizationsAlibaba GroupMax Planck Institute for InformaticsPeking UniversityThe Chinese University of Hong KongUniversity of Toronto

Why you should read this

Presents RBGNet, a point-based 3D object detection framework that improves bounding box prediction on point clouds by combining foreground-biased point sampling with ray-based feature grouping to model object surface geometry.

As a fundamental problem in computer vision, 3D object detection is experiencing rapid growth. To extract the point-wise features from the irregularly and sparsely distributed points, previous methods usually take a feature grouping module to aggregate the point features to an object candidate. However, these methods have not yet leveraged the surface geometry of foreground objects to enhance grouping and 3D box generation. In this paper, we propose the RBGNet framework, a voting-based 3D detector for accurate 3D object detection from point clouds. In order to learn better representations of object shape to enhance cluster features for predicting 3D boxes, we propose a ray-based feature grouping module, which aggregates the point-wise features on object surfaces using a group of determined rays uniformly emitted from cluster centers. Considering the fact that foreground points are more meaningful for box estimation, we design a novel foreground biased sampling strategy in downsample process to sample more points on object surfaces and further boost the detection performance. Our model achieves state-of-the-art 3D detection performance on ScanNet V2 and SUN RGB-D with remarkable performance gains. Code will be available at https://github.com/Haiyang-W/RBGNet.

Added

2026-09-26

Cell Detection with Star-convex Polygons

Cell Detection with Star-convex Polygons

Uwe Schmidt, Martin Weigert, Coleman Broaddus, Gene Myers

OrganizationsCenter for Systems Biology DresdenMax Planck Institute of Molecular Cell Biology and GeneticsTechnische Universität Dresden

Why you should read this

Proposes using star-convex polygons within convolutional neural networks to accurately detect and segment crowded cell nuclei in microscopy images, overcoming the merging and suppression errors of bounding-box approaches without requiring shape refinement.

Automatic detection and segmentation of cells and nuclei in microscopy images is important for many biological applications. Recent successful learning-based approaches include per-pixel cell segmentation with subsequent pixel grouping, or localization of bounding boxes with subsequent shape refinement. In situations of crowded cells, these can be prone to segmentation errors, such as falsely merging bordering cells or suppressing valid cell instances due to the poor approximation with bounding boxes. To overcome these issues, we propose to localize cell nuclei via star-convex polygons, which are a much better shape representation as compared to bounding boxes and thus do not need shape refinement. To that end, we train a convolutional neural network that predicts for every pixel a polygon for the cell instance at that position. We demonstrate the merits of our approach on two synthetic datasets and one challenging dataset of diverse fluorescence microscopy images.

Added

2026-09-24

Multi-view Convolutional Neural Networks for 3D Shape Recognition

Multi-view Convolutional Neural Networks for 3D Shape Recognition

Hang Su, Subhransu Maji, Evangelos Kalogerakis, Erik Learned-Miller

OrganizationsUniversity of Massachusetts Amherst

Why you should read this

Introduces a multi-view convolutional neural network architecture that pools standard 2D rendered views into a compact descriptor, demonstrating that 2D image representations outperform native 3D voxel and mesh models on 3D shape and sketch recognition.

A longstanding question in computer vision concerns the representation of 3D shapes for recognition: should 3D shapes be represented with descriptors operating on their native 3D formats, such as voxel grid or polygon mesh, or can they be effectively represented with view-based descriptors? We address this question in the context of learning to recognize 3D shapes from a collection of their rendered views on 2D images. We first present a standard CNN architecture trained to recognize the shapes' rendered views independently of each other, and show that a 3D shape can be recognized even from a single view at an accuracy far higher than using state-of-the-art 3D shape descriptors. Recognition rates further increase when multiple views of the shapes are provided. In addition, we present a novel CNN architecture that combines information from multiple views of a 3D shape into a single and compact shape descriptor offering even better recognition performance. The same architecture can be applied to accurately recognize human hand-drawn sketches of shapes. We conclude that a collection of 2D views can be highly informative for 3D shape recognition and is amenable to emerging CNN architectures and their derivatives.

Added

2026-09-11

Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains

Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains

Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, Ren Ng

OrganizationsGoogleUniversity of California BerkeleyUniversity of California, San Diego

Why you should read this

Demonstrates how Fourier feature mappings overcome the inherent spectral bias of coordinate-based neural networks, enabling them to fit high-frequency details in low-dimensional computer vision and graphics tasks.

We show that passing input points through a simple Fourier feature mapping enables a multilayer perceptron (MLP) to learn high-frequency functions in low-dimensional problem domains. These results shed light on recent advances in computer vision and graphics that achieve state-of-the-art results by using MLPs to represent complex 3D objects and scenes. Using tools from the neural tangent kernel (NTK) literature, we show that a standard MLP fails to learn high frequencies both in theory and in practice. To overcome this spectral bias, we use a Fourier feature mapping to transform the effective NTK into a stationary kernel with a tunable bandwidth. We suggest an approach for selecting problem-specific Fourier features that greatly improves the performance of MLPs for low-dimensional regression tasks relevant to the computer vision and graphics communities.

Added

2026-09-11

License

Published with permission

DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation

DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation

Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, Steven Lovegrove

OrganizationsMassachusetts Institute of TechnologyMetaUniversity of Washington

Why you should read this

Introduces a continuous neural implicit representation that learns signed distance functions across entire shape classes, enabling high-quality 3D shape reconstruction, interpolation, and completion from partial inputs with an order-of-magnitude reduction in model size.

Computer graphics, 3D computer vision and robotics communities have produced multiple approaches to representing 3D geometry for rendering and reconstruction. These provide trade-offs across fidelity, efficiency and compression capabilities. In this work, we introduce DeepSDF, a learned continuous Signed Distance Function (SDF) representation of a class of shapes that enables high quality shape representation, interpolation and completion from partial and noisy 3D input data. DeepSDF, like its classical counterpart, represents a shape's surface by a continuous volumetric field: the magnitude of a point in the field represents the distance to the surface boundary and the sign indicates whether the region is inside (-) or outside (+) of the shape, hence our representation implicitly encodes a shape's boundary as the zero-level-set of the learned function while explicitly representing the classification of space as being part of the shapes interior or not. While classical SDF's both in analytical or discretized voxel form typically represent the surface of a single shape, DeepSDF can represent an entire class of shapes. Furthermore, we show state-of-the-art performance for learned 3D shape representation and completion while reducing the model size by an order of magnitude compared with previous work.

Added

2026-09-10

Structured 3D Latents for Scalable and Versatile 3D Generation

Structured 3D Latents for Scalable and Versatile 3D Generation

Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, Jiaolong Yang

OrganizationsMicrosoftTsinghua UniversityUniversity of Science and Technology of China

Why you should read this

Establishes a unified Structured Latent (SLAT) representation that enables high-quality, versatile 3D generation across Radiance Fields, 3D Gaussians, and meshes, outperforming existing methods and introducing flexible output and editing capabilities.

We introduce a novel 3D generation method for versatile and high-quality 3D asset creation. The cornerstone is a unified Structured LATent (SLAT) representation which allows decoding to different output formats, such as Radiance Fields, 3D Gaussians, and meshes. This is achieved by integrating a sparsely-populated 3D grid with dense multiview visual features extracted from a powerful vision foundation model, comprehensively capturing both structural (geometry) and textural (appearance) information while maintaining flexibility during decoding. We employ rectified flow transformers tailored for SLAT as our 3D generation models and train models with up to 2 billion parameters on a large 3D asset dataset of 500K diverse objects. Our model generates high-quality results with text or image conditions, significantly surpassing existing methods, including recent ones at similar scales. We showcase flexible output format selection and local 3D editing capabilities which were not offered by previous models. Code, model, and data will be released.

Added

2026-05-19

License

Published with permission