Toggle Main Menu Toggle Search

Open Access padlockePrints

Bridging the semantic gap for zero-shot object counting with language-augmented pseudo-exemplars

Lookup NU author(s): Dr Shidong WangORCiD

Downloads

Full text for this publication is not currently held within this repository. Alternative links are provided below where available.


Abstract

© 2026 Elsevier B.V.Zero-shot Object Counting (ZSC) aims to estimate the number of objects in novel categories solely based on class names, without requiring any annotated examples. Although recent advances in vision-language models, such as CLIP, have significantly improved performance, existing approaches still suffer from a persistent semantic gap between global text embeddings and local visual representations, leading to biased or inconsistent counting in complex scenes. This discrepancy often results from the inability of coarse text features to capture fine-grained visual distinctions and contextual nuances. To address these limitations, we propose Language-augmented Pseudo-exemplars (LAPE), a novel and comprehensive ZSC framework that performs both coarse- and fine-grained cross-modal alignment through a unified and trainable architecture. Specifically, we first leverage a large language model to enrich textual semantics, generating diverse and context-aware descriptions that extend beyond simple class names. The query image features and these enhanced text features are then aligned across modalities. Subsequently, a sparse semantic projection mechanism maps the global text embedding onto a compact set of the most relevant visual features, synthesizing language-augmented pseudo-exemplar representations. These pseudo-exemplars, together with the enriched textual features, jointly guide a density regression head for precise spatial distribution estimation, effectively suppressing background noise and enhancing focus on discriminative regions. Extensive experiments on FSC-147 and cross-dataset benchmarks demonstrate that our approach substantially strengthens text-vision alignment and achieves state-of-the-art counting accuracy, showcasing strong generalization capability to unseen categories even in complex scenarios. Code is available at https://github.com/WJWu20/LAPE .


Publication metadata

Author(s): Wu W, Ren Y, Wang S, Zhang H

Publication type: Article

Publication status: Published

Journal: Pattern Recognition Letters

Year: 2026

Volume: 208

Pages: 86-92

Online publication date: 23/07/2026

Acceptance date: 22/07/2026

ISSN (print): 0167-8655

ISSN (electronic): 1872-7344

Publisher: Elsevier

URL: https://doi.org/10.1016/j.patrec.2026.07.026

DOI: 10.1016/j.patrec.2026.07.026


Altmetrics

Altmetrics provided by Altmetric


Share