🐠">
[go: up one dir, main page]

Xinyu Huang (黄新宇)

I am a researcher at ByteDance Seed, working on multimodal generation.

I obtained my Ph.D. from Fudan University, advised by Prof. Rui Feng. I was also a visiting student at MMLab, Nanyang Technological University, advised by Prof. Ziwei Liu.

From 2020 to 2024, I worked on visual perception and created the Recognize Anything Model (RAM) family, a series of open-source image perception models that exceed OpenAI's CLIP by more than 20 points in fine-grained perception. This work was done at OPPO Research Institute and IDEA, together with Youcai Zhang, Prof. Yandong Guo, and Prof. Lei Zhang.

From 2024 to 2025, I co-led the development (pre-training, SFT, and RL) of large multimodal models at TikTok AI Innovation Center.

Xinyu Huang

Research & Work

Seedream 5.0 Pro

Seedream 5.0 Pro

ByteDance Seed, 2026

project page / tech blog

Seedream 5.0 Pro is a multimodal image creation model that reasons before it draws. It plans logical layouts for text-dense infographics, grounds point, lasso, sketch, and color signals for pixel-level interactive editing, decomposes an image into 10+ editable transparent layers, and natively renders text in over ten languages.

MGPO

Multi-Turn Grounding-Based RL (MGPO)

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

ACL 2026

arXiv / code

MGPO lets LMMs iteratively focus on key image regions through automatic grounding, achieving superior performance on high-resolution visual tasks without any grounding annotations.

RAM++

Recognize Anything Plus Model (RAM++)

Open-Set Image Tagging with Multi-Grained Text Supervision

CVPR 2024, Multimodal Foundation Models Workshop

arXiv / code

RAM++ is the next generation of RAM, which recognizes any category with high accuracy, covering both predefined common categories and diverse open-set categories.

RAM

Recognize Anything Model (RAM)

Recognize Anything: A Strong Image Tagging Model

CVPR 2024 Workshop on Multimodal Foundation Models

project page / arXiv / demo / code

RAM is an image tagging model that recognizes any common category with high accuracy.

Tag2Text

Tag2Text Vision-Language Model

Tag2Text: Guiding Vision-Language Model via Image Tagging

ICLR 2024

project page / arXiv / demo / code

Tag2Text is a vision-language model guided by tagging, which supports tagging and comprehensive captioning simultaneously.

IDEA

IDEA

Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training

ACM MM 2022

arXiv / code

IDEA provides more explicit textual supervision for visual models, including valuable tags and texts composed of multiple tags.

Open-Source Projects

Recognize Anything family

Recognize Anything Family

The Recognize Anything Model (RAM) family, demonstrating superior image recognition ability across common and open-set categories.

RAM-Grounded-SAM

RAM-Grounded-SAM

The RAM family married with Grounded-SAM, which can automatically recognize, detect, and segment anything in an image.

Academic Service

Reviewer for CVPR, ICCV, ECCV, ICLR, NeurIPS, IJCV, WACV, and ACM MM.