Jinhong Deng

Researcher at Meituan LongCat Team

I am currently a Researcher at Meituan LongCat Team (Beidou Program). I received my Ph.D. degree from the School of Computer Science and Engineering, University of Electronic Science and Technology of China (SCSE, UESTC), supervised by Prof. Lixin Duan and Prof. Wen Li. Before that, I obtained the B.Eng degree from the School of Information Science and Technology, Southwest Jiaotong University (SIST, SWJTU) in 2019. During my Ph.D., I was also honored to work with Joey Tianyi Zhou and Yang He as a visiting student at A*STAR Centre for Frontier AI Research (CFAR), Singapore. I sincerely appreciate everyone I have encountered and every moment of joy and challenge along the way.

Email: jhdengvision AT gmail.com, [Google Scholar], [GitHub].

During my Ph.D., my research mainly focused on computer vision and transfer learning, especially visual perception tasks such as object detection and semantic segmentation, aiming to build a robust visual system capable of effectively operating in dynamic, open environments. My works cover cross-domain object detection (UMT, CVPR 2021, 500+ citations; HT, CVPR 2023, 100+ citations), semantic segmentation (BAPA-Net, ICCV 2021 Oral), open-vocabulary detection (ROOT, TIP 2026), and source-free domain adaptation (DMCD, AAAI 2022 Oral; Balanced Teacher, TCSVT 2024). I was also interested in model efficiency and proposed visual token compression for multimodal LLMs (SCOPE, NeurIPS 2025).

Currently, I mainly work on pre-train of multimodal/omnimodal foundation models, specifically focusing on topics such as audio-video captioning, understanding, and interaction, as well as long context and memory. If you are interested in my work, please feel free to contact me. I am also open to hosting interns and can provide rich guidance and ample computing resources.

My long-term goal is to build AGI for the digital world and ultimately extend it to the physical world. I am particularly interested in enabling intelligent agents to perceive, remember, and interact—continually understanding evolving multimodal environments and collaborating naturally with humans over time.


Updates


Publications

HoloCount: A Holistic Visual Counting Benchmark for MLLMs
Jinhong Deng, Limeng Qiao, Guanglu Wan.
A holistic visual counting benchmark to evaluate the counting ability of multimodal large language models.
arXiv preprint, 2026.
Mapping Text to Multiplex Graph: Prompt Compression as Lévy Walk-Guided Graph Pruning
Yaxin Gao, Yao Lu, Jinhong Deng, Jiaqi Nie, Zhe Tang, Jian Zhang, Zhaowei Zhu, Shanqing Yu, Qi Xuan, Joey Tianyi Zhou.
Compresses long prompts by modeling text as a multiplex graph and pruning it via Lévy walk-guided graph pruning.
Knowledge-Based Systems, 2026.
[paper]
Instance-Free Domain Adaptive Object Detection
Hengfu Yu, Jinhong Deng, Lixin Duan, Wen Li.
Adapts object detectors across domains without requiring any instance-level information.
arXiv preprint, 2026.
[paper]
SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodal LLMs
Jinhong Deng, Wen Li, Joey Tianyi Zhou, Yang He
Prunes redundant visual tokens in multimodal LLMs with a saliency-coverage oriented strategy for efficient inference.
NeurIPS 2025.
[paper][code]
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference
Yuhang Yang*, Jinhong Deng*, Wen Li, Lixin Duan. * equal contribution.
A residual attention mechanism for training-free dense vision-language inference.
CVPR 2025.
[paper][code]
Towards Unsupervised Model Selection for Domain Adaptive Object Detection
Hengfu Yu*, Jinhong Deng*, Wen Li, Lixin Duan. * equal contribution.
An unsupervised criterion to select the best-performing model for domain adaptive object detection.
Neural Information Processing Systems (NeurIPS), 2024.
[paper][code]
Balanced Teacher for Source-free Object Detection
Jinhong Deng, Wen Li, Lixin Duan.
A balanced teacher framework for source-free object detection without accessing any source data.
IEEE Transactions on Circuits and Systems for Video Technology, 2024.
[paper]
Domain Adaptive Detection of MAVs: A Benchmark and Noise Suppression Network
Yin Zhang, Jinhong Deng, Peidong Liu, Wen Li, Shiyu Zhao.
A new benchmark and a noise suppression network for domain-adaptive micro aerial vehicle detection.
IEEE Transactions on Automation Science and Engineering, 2024.
[paper]
Cross-domain Detection Transformer based on Spatial-aware and Semantic-aware Token Alignment
Jinhong Deng, Xiaoyue Zhang, Wen Li, Lixin Duan, Dong Xu.
A DETR-based cross-domain detector with spatial-aware and semantic-aware token alignment.
IEEE Transactions on Multimedia, 2023.
[paper]
Harmonious Teacher for Cross-domain Object Detection
CVPR 2023
Jinhong Deng, Dongli Xu, Wen Li, Lixin Duan.
A harmonious teacher that decouples cross-domain detection into separate classification and localization stages.
[paper] [code]
Motion and Appearance Adaptation for Cross-Domain Motion Transfer
ECCV 2022
Borun Xu, Biao Wang, Jinhong Deng, Jiale Tao, Tiezheng Ge, Yuning Jiang, Wen Li, Lixin Duan.
Adapts both motion and appearance for cross-domain motion transfer.
[paper]
Undoing the Damage of Label Shift for Cross-domain Semantic Segmentation
CVPR 2022
Yahao Liu, Jinhong Deng, Jiale Tao, Tong Chu, Lixin Duan, Wen Li.
Mitigates the label shift between domains for cross-domain semantic segmentation.
[paper] [code]
Revisiting AP Loss for Dense Object Detection: Adaptive Ranking Pair Selection
CVPR 2022
Dongli Xu, Jinhong Deng, Wen Li.
A revised AP loss with adaptive ranking pair selection for dense object detection.
[paper] [code]
Denoised Maximum Classifier Discrepancy for Source-Free Unsupervised Domain Adaptation
AAAI 2022
Tongchu, Yahao Liu, Jinhong Deng, Wen Li, Lixin Duan.
Accepted as oral paper

Denoises unreliable target samples to make classifier discrepancy robust for source-free unsupervised domain adaptation.
[paper]
BAPA-Net: Boundary Adaptation and Prototype Alignment for Cross-domain Semantic Segmentation
ICCV 2021
Yahao Liu, Jinhong Deng, Xinchen Gao, Wen Li, Lixin Duan.
Accepted as oral paper

Aligns source and target class prototypes and adapts the decision boundary to reduce the domain gap in semantic segmentation.
[paper] [code]
Unbiased Mean Teacher for Cross-domain Object Detection
CVPR 2021
Jinhong Deng, Wen Li, Yuhua Chen, Lixin Duan. CVPR 2021
We propose an unbiased mean teacher method to improve the cross-domain ability of Faster RCNN.
[paper] [code]

Experiences

2026.01—Now: Algorithm Researcher at Meituan LongCat Team

2021.09—2025.12: Ph.D. Student at SCSE, UESTC

2024.09—2025.09: Visiting Ph.D. Student, A*STAR Centre for Frontier AI Research (CFAR)

2019.09—2021.05: Master's Student at SCSE, UESTC

2015.09—2019.05: Undergraduate Student at Southwest Jiaotong University


Honors & Awards

2025, Outstanding Doctoral Dissertation of UESTC.
2025, Outstanding Graduate of Sichuan Province.
2025, Outstanding Graduate of UESTC.
2024, UESTC Academic Rising Star Award, UESTC.
2021, UESTC Academic Promising Talent Award, UESTC.
2019, The First Prize Graduate Academic Scholarships, UESTC.
2019, Outstanding Graduate from Southwest Jiaotong University.
2019, TOP 10, SCADA Data Missing Repair Competition.
2019, TOP 3, AI Challenger 2018 in Weather Forecasting.
2017, National First Prize, 3Chuang: Innovation, Creativity, Entrepreneurship Competition.
2017, Third Prize, ACM Competition of Southwest Jiaotong University.
2016, 2017, National Encouragement scholarship, SWJTU.

Professional Activities


Skills


Misc