Sketch-based manga retrieval using manga109 dataset
Yusuke MatsuiKota ItoYuji AramakiAzuma FujimotoToru OgawaToshihiko YamasakiKiyoharu Aizawa
Introduces Manga109 alongside a real-time, sketch-based retrieval framework that enables fast sub-image comic search through screentone-invariant edge features and interactive reranking.
As digital comic platforms grow, users face significant difficulty discovering content because commercial archives rely almost exclusively on basic text searches by title or author. This keyword-only approach fails to search visual content, which is problematic for Japanese comics (manga) that are defined by distinctive line art, non-textured graphics, and complex multi-frame page layouts. The article addresses this limitation by developing and evaluating an efficient, content-based retrieval system that enables users to search manga collections using hand-drawn sketches.
The system workflow preprocesses manga pages by identifying and filtering out blank inter-frame margins, then applies an object detector to locate candidate visual regions. To improve line detection, screentone shading patterns are removed before extracting edge orientation histogram features, which capture the geometric outlines of drawings regardless of scale. These visual features are compressed into compact binary codes using product quantization, enabling fast approximate nearest-neighbor matching against sketch queries. To support realistic testing, the researchers constructed Manga109, a publicly available benchmark dataset comprising 109 titles and 21,142 pages created by 94 professional artists.
Evaluation results show substantial improvements in both retrieval accuracy and processing efficiency. Across comparative tests, the proposed feature representation consistently outperformed standard image retrieval methods, achieving higher recall rates across varying visual targets. Feature compression enabled rapid querying, allowing the system to search across 14 million candidate regions from 21,142 pages in 70 milliseconds using 204 megabytes of memory on a single parallelized computer. Furthermore, interactive reranking methods, including relevance feedback (using retrieved images as new queries) and query retouching (editing existing sketches), allowed users to quickly refine searches to locate specific characters or visual traits across diverse titles.
These findings demonstrate that sketch-based retrieval is commercially viable at scale without requiring expensive computational infrastructure. Because the pipeline relies on compact indexing and lightweight vector matching, digital storefronts and archives can implement visual search on standard server hardware with minimal operational overhead. This offers an intuitive exploration mechanism for consumers while providing a standardized, legally cleared dataset to accelerate academic research in comic processing.
To build upon this work, developers should integrate text-filtering techniques to prevent Japanese text and speech balloons from appearing as false-positive sketch matches. Combining sketch queries with standard keyword metadata would also provide a hybrid search experience. While initial localization results indicate that finding small objects across millions of candidates remains challenging and subject to user drawing variability, the underlying architecture provides high confidence for large-scale, low-latency visual retrieval.
- Paper: Aggregating Local Image Descriptors into Compact Codes, Hervé Jégou et al. (2012). It provides essential foundations for large-scale image indexing, descriptor aggregation, and product quantization used for scalable approximate nearest-neighbor search.
- Paper: Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories, Svetlana Lazebnik et al. (2006). It establishes spatial pyramid matching and edge orientation histogram representations that underlie spatial feature descriptions in visual retrieval.
- Paper: Iterative Quantization: A Procrustean Approach to Learning Binary Codes for Large-Scale Image Retrieval, Yunchao Gong et al. (2013). It explores compact binary code learning and quantization strategies for efficient similarity search across large image collections.
- Paper: Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search, Hervé Jégou et al. (2008). It details indexing and geometric verification techniques in inverted file systems that inform fast visual search pipelines.
- Paper: Content-Based Image Retrieval at the End of the Early Years, Arnold W.M. Smeulders et al. (2000). It surveys foundational concepts, feature extraction principles, and interactive relevance feedback in content-based image retrieval.
- Paper: Shape Matching and Object Recognition Using Shape Contexts, Serge Belongie et al. (2002). It presents foundational shape-context and contour-matching methodologies critical for sketch-to-image alignment and retrieval.
- Paper: Scalable Nearest Neighbor Algorithms for High Dimensional Data, Marius Muja et al. (2014). It introduces scalable approximate nearest-neighbor matching algorithms that underpin high-dimensional feature retrieval systems.
- Paper: The use of MMR, diversity-based reranking for reordering documents and producing summaries, Jaime Carbonell et al. (1998). It defines classic interactive reranking and diversity principles that motivate sketch-based query refinement and feedback mechanisms.
- Paper: Fine-Tuning CNN Image Retrieval with No Human Annotation, Filip Radenovic et al. (2017). It extends visual retrieval pipelines by training convolutional neural networks with generalized pooling and metric learning to replace hand-crafted edge descriptors.
- Paper: Accelerating Large-Scale Inference with Anisotropic Vector Quantization, Ruiqi Guo et al. (2020). It advances vector quantization by introducing anisotropic loss for faster and more accurate nearest-neighbor search over high-dimensional image embeddings.
- Paper: DiskANN: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node, Suhas Jayaram Subramanya et al. (2019). It scales approximate nearest-neighbor indexing beyond in-memory product quantization to billion-point datasets on single-node SSD architectures.
- Paper: Image-to-Image Translation with Conditional Adversarial Networks, Phillip Isola et al. (2017). It demonstrates conditional adversarial translation frameworks that advance cross-domain sketch-to-image synthesis and matching.
