Wenqing Wang
I am Wenqing Wang (ηζι), a Computer Engineering PhD candidate at Northeastern University. I am advised by Prof. Yun Raymond Fu (Member of the Academy of Europe, European Academy of Sciences and Arts; Fellow of AAAS, IEEE, AAAI, ACM, UAI), and I am a member of the SMILE Lab. I earned my M.S. in Computer Science from Northeastern University in August 2025, and my B.S. in Computer Science and B.A. in Psychology from the University of California, Davis.
My research interests span computer vision, generative AI, and
general AI. In particular, I focus on
generative AI [world modeling, image/video generation & understanding, agents, physical AI, diffusion, VLM/VLA/LLM];
neural rendering [NeRF, Gaussian Splatting, 3D novel-view synthesis]; and
general AI [ML, efficiency, brain-inspired algorithms, AI-driven longevity/health, AI-cosmology, autonomous driving].
Actively seeking internships/full-time jobs.
News
September, 2026 09/26 | RacerGS: Residual-Autoregressive Context for 3D Gaussian Splatting Compression is accepted to NeurIPS 2026 π |
|---|---|
September, 2026 09/26 | Started research scientist internship at Adobe π |
June, 2026 06/26 | Started research assistant role at NEC π |
|
May, 2026
05/26
|
Started research scientist internship at NVIDIA ππ |
|
April, 2026
04/2026
|
Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis is accepted to ICML 2026 π |
|
Sep 21, 2025
09/21/25
|
BEAR: Benchmarking Multimodal Language Models for Atomic Embodied Reasoning Abilities is accepted to NeurIPS 2025 LLM Evaluation Workshop π |
|
September 15, 2025
09/15/25
|
Started research scientist internship at Samsung π |
|
Jul 11, 2025
07/11/25
|
RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis is accepted to ICCV 2025 Workshop on Artificial Social Intelligence π |
|
June 2, 2025
06/02/25
|
Started research scientist internship at Socure π |
|
Dec 7, 2024
12/07/24
|
Preprint of Text-to-3D Gaussian Splatting with Physics-Grounded Motion Generation is available |
|
Oct 7, 2024
10/07/24
|
EmoGene: Audio-Driven Emotional 3D Talking-Head Generation is accepted to FG 2025 π |
|
May 24, 2024
05/24/24
|
Started research scientist internship at bitHuman π |
|
Sep 1, 2022
09/01/22
|
Started Ph.D. at Northeastern University |
|
Jun 10, 2022
06/10/22
|
Graduated from University of California, Davis |
Publications
Clean Forcing: Drift-Resistant Autoregressive Video Diffusion with a Frozen Base.Under Review
CuckooCache: A Training-Free KV Cache via Cuckoo Hashing and Locality-Sensitive Routing.Under Review
RacerGS: Residual-Autoregressive Context for 3D Gaussian Splatting Compression.NeurIPS, 2026
LatentPhyGS: Dynamic Physics-Grounded 3DGS with Material-Aware Latents.Under Review
Research Projects
Physics-Aware World Model
- Fine-tuned Wan2.1-T2V-1.3B with Diffusers and PEFT LoRA to predict sixteen future frames from an initial scene and object motion conditions.
Interactive Game World Model
- Built live editing for OmniDreams using in-place cross-attention KV replacement and ReCache, allowing prompt changes at run time, such as weather swaps, while preserving scene history.
Agentic Adaptive Benchmarking Pipeline for World Models
- Built a closed-loop agentic evaluation pipeline connecting an agent controller, an HDMap scenario generator, OmniDreams, and a Qwen3-based VLA judge.
Efficient, High-Quality Video Editing and Generation with Trajectory Conditioning
- Designed a lightweight reference-conditioning method for editing selected video objects while preserving their identities and the surrounding content.
Multimodal Adapter for Improved Image Understanding and Generation
- Connected Qwen2.5-VL to FLUX.1-schnell through a lightweight BLIP-2 Q-Former and trainable projection layers, bridging language and diffusion representations.
Query-Aware KV Cache Compression for Efficient Inference
- Extended Hugging Face DynamicCache to assign separate KV retention budgets to text, image, and video tokens using attention scores, query similarity, and recency.
Fun Projects
Experience
-
Sep. 2026 β Present Research Scientist Intern
Adobe
San Jose, CA- Enabled personalized, photorealistic style transfer for Adobe Lightroom users from only a few reference images.
- Developed and trained a style encoder that distills the images into compact style tokens.
-
Jun. 2026 β Aug. 2026 Research Assistant
NEC
Princeton, NJ (Remote)- Designed a compact scene representation that distills multimodal (RGB + LiDAR) Gaussian Splatting into a single differentiable signed distance field that returns geometry, semantics, physical properties, and calibrated uncertainty from a single query.
- Validated the field distillation on the FOCI benchmark.
-
May 2026 β Aug. 2026 Research Scientist Intern
NVIDIA
Santa Clara, CA- Developed Clean Forcing, an efficient method for reducing drift and improving long-horizon quality in autoregressive video and world models; filed a patent and integrated it into FlashDreams.
- Enhanced the interactability of OmniDreams, a Cosmos Predict 2.5-based autonomous driving world model, with prompt hot-swapping, restyle LoRAs, Clean Forcing, and conditioned object insertion.
-
Sep. 2025 β Dec. 2025 Research Scientist Intern
Samsung
Mountain View, CA- Improved spatial and temporal consistency in camera- and motion-controllable video generation built on Diffusion Forcing Transformers (DFoT) with history guidance.
- Designed and evaluated new architectures for DFoT-based video diffusion models.
-
Jun. 2025 β Aug. 2025 Research Scientist Intern
Socure
Boston, MA- Post-trained the Qwen-Image-Edit model to generate synthetic ID images, including text and profile photos, to enhance ID fraud detection.
- Trained an ID fraud detection model using synthetic ID images and improved its robustness with a StyleGAN3-based evaluator.
-
May 2024 β Aug. 2024 Research Scientist Intern
bitHuman
Boston, MA- Implemented an audio-driven 3D emotional talking-head generation framework.
- Applied a variational autoencoder to generate dynamic facial landmarks driven by input audio.
-
Sep. 2023 β May 2024 Research Assistant
Northeastern University, Synergetic Media Learning Lab
Boston, MA- Developed a text-to-3D Gaussian Splatting pipeline to simulate physics-grounded motion across diverse materials and object properties.
- Integrated an LLM with evolutionary search and chain-of-thought (CoT) to iteratively optimize prompts, generating more precise 3D objects.
-
Sep. 2022 β Sep. 2023 Research Assistant
Northeastern University, Cognitive Embodied Social Agents Research Lab
Boston, MA- Formulated hierarchical deep Bayesian reinforcement learning methods for Partially Observable Markov Decision Processes (POMDPs).
- Integrated Theory of Mind into the POMDP framework to enhance decision-making processes.
-
Jun. 2021 β Jun. 2022 Research Assistant
University of California, Davis, FoxLab
Davis, CA- Applied a Mask R-CNN model to detect 2D keypoints for tracking primates (Rhesus macaques).
- Enabled the model to recognize primate behaviors through analysis of gestures and body features.
Education
-
Sep. 2022 - Present -
Sep. 2022 - Aug. 2025 -
Sep. 2018 - Jun. 2022
Grants
Modal Neurips $10,000 GPU Grant, 2026
Modal ICLR $10,000 GPU Grant, 2026
ICCV Travel Grant, 2025
FG PhD Doctoral Consortium Travel Grant, 2025
Teaching
Courses TA'd: CS5100 Foundations of Artificial Intelligence; CS5150/4150 Game Artificial Intelligence; CS5340: Computer/Human Interaction; CS3200: Intro to Database Systems
Academic Services
Conference Reviewer: CVPR, NeurIPS, ICLR, ICML, ICCV, ECCV, AAAI, FG, MM
Journal Reviewer: Neural Computing and Applications, TPAMI, TKDD, Neural Networks, IEEE TNNLS, KnowledgeβBased Systems, AI Reviews, Multimedia Tools and Applications