Wenqing Wang

images/prof_pic.png
images/prof_pic.png

I am Wenqing Wang (ηŽ‹ζ–‡ι’), a Computer Engineering PhD candidate at Northeastern University. I am advised by Prof. Yun Raymond Fu (Member of the Academy of Europe, European Academy of Sciences and Arts; Fellow of AAAS, IEEE, AAAI, ACM, UAI), and I am a member of the SMILE Lab. I earned my M.S. in Computer Science from Northeastern University in August 2025, and my B.S. in Computer Science and B.A. in Psychology from the University of California, Davis.

My research interests span computer vision, generative AI, and general AI. In particular, I focus on generative AI [world modeling, image/video generation & understanding, agents, physical AI, diffusion, VLM/VLA/LLM]; neural rendering [NeRF, Gaussian Splatting, 3D novel-view synthesis]; and general AI [ML, efficiency, brain-inspired algorithms, AI-driven longevity/health, AI-cosmology, autonomous driving].

Actively seeking internships/full-time jobs.


News

September, 2026   
09/26  
RacerGS: Residual-Autoregressive Context for 3D Gaussian Splatting Compression is accepted to NeurIPS 2026 πŸŽ‰
September, 2026   
09/26  
Started research scientist internship at Adobe πŸš€
June, 2026   
06/26  
Started research assistant role at NEC πŸš€
May, 2026   
05/26  
Started research scientist internship at NVIDIA πŸš€πŸ’š
April, 2026   
04/2026  
Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis is accepted to ICML 2026 πŸŽ‰
Sep 21, 2025   
09/21/25  
BEAR: Benchmarking Multimodal Language Models for Atomic Embodied Reasoning Abilities is accepted to NeurIPS 2025 LLM Evaluation Workshop πŸŽ‰
September 15, 2025   
09/15/25  
Started research scientist internship at Samsung πŸš€
Jul 11, 2025   
07/11/25  
RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis is accepted to ICCV 2025 Workshop on Artificial Social Intelligence πŸŽ‰
June 2, 2025   
06/02/25  
Started research scientist internship at Socure πŸš€
Dec 7, 2024   
12/07/24  
Preprint of Text-to-3D Gaussian Splatting with Physics-Grounded Motion Generation is available
Oct 7, 2024   
10/07/24  
EmoGene: Audio-Driven Emotional 3D Talking-Head Generation is accepted to FG 2025 πŸŽ‰
May 24, 2024   
05/24/24  
Started research scientist internship at bitHuman πŸš€
Sep 1, 2022   
09/01/22  
Started Ph.D. at Northeastern University
Jun 10, 2022   
06/10/22  
Graduated from University of California, Davis

Publications

  1.  

    Clean Forcing preview
    Clean Forcing: Drift-Resistant Autoregressive Video Diffusion with a Frozen Base.
    Wenqing Wang et al.
    Under Review
  2.  

    CuckooCache preview
    CuckooCache: A Training-Free KV Cache via Cuckoo Hashing and Locality-Sensitive Routing.
    Wenqing Wang and Yun Fu
    Under Review
  3.  

    RacerGS preview
    RacerGS: Residual-Autoregressive Context for 3D Gaussian Splatting Compression.
    Wenqing Wang, Huimin Zeng, Yun Fu
    NeurIPS, 2026
  4.  

    LatentPhyGS preview
    LatentPhyGS: Dynamic Physics-Grounded 3DGS with Material-Aware Latents.
    Wenqing Wang and Yun Fu
    Under Review
  5.  

    images/publication_preview/bear.png
    Dissecting Embodied Abilities in Multimodal Language Models through Skill-Level Evaluation and Diagnosis
    Yu Qi, Haibo Zhao, Ziyu Guo, Siyuan Ma, Ziyan Chen, Yaokun Han, Renrui Zhang, Zitiantao Lin, Shiji Xin, Yijian Huang, Boce Hu, Kai Cheng, Jiayi Zhang, Peiheng Wang, Jiazheng Liu, Wenqing Wang, Yiran Qin, Haojie Huang, Lawson L. S. Wong
    ICML, 2026
  6.  

    images/publication_preview/wang2024audiodrivenemotional3dtalkinghead.gif
    EmoGene: Audio-Driven Emotional 3D Talking-Head Generation
    Wenqing Wang, Yun Fu
    2025 IEEE 19th International Conference on Automatic Face and Gesture Recognition (FG)
  7.  

    images/publication_preview/realtalk.gif
    RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis
    Wenqing Wang, Yun Fu
    ICCV 2025 Workshop on Artificial Social Intelligence, 2025
  8.  

    images/publication_preview/bear_w.png
    BEAR: Benchmarking Multimodal Language Models for Atomic Embodied Reasoning Abilities
    Yu Qi, Haibo Zhao, Ziyu Guo, Siyuan Ma, Ziyan Chen, Yaokun Han, Renrui Zhang, Zitiantao Lin, Shiji Xin, Yijian Huang, Kai Cheng, Peiheng Wang, Jiazheng Liu, Jiayi Zhang, Yizhe Zhu, Wenqing Wang, Yiran Qin, Xupeng Zhu, Haojie Huang, Lawson L.S. Wong
    NeurIPS 2025 LLM Evaluation Workshop, 2025
  9.  

    images/publication_preview/wang2024textto3dgaussiansplattingphysicsgrounded.gif
    Text-to-3D Gaussian Splatting with Physics-Grounded Motion Generation
    Wenqing Wang, Yun Fu
    Preprint, 2024

Research Projects

  • Physics-Aware World Model
    • Fine-tuned Wan2.1-T2V-1.3B with Diffusers and PEFT LoRA to predict sixteen future frames from an initial scene and object motion conditions.
  • Interactive Game World Model
    • Built live editing for OmniDreams using in-place cross-attention KV replacement and ReCache, allowing prompt changes at run time, such as weather swaps, while preserving scene history.
  • Agentic Adaptive Benchmarking Pipeline for World Models
    • Built a closed-loop agentic evaluation pipeline connecting an agent controller, an HDMap scenario generator, OmniDreams, and a Qwen3-based VLA judge.
  • Efficient, High-Quality Video Editing and Generation with Trajectory Conditioning
    • Designed a lightweight reference-conditioning method for editing selected video objects while preserving their identities and the surrounding content.
  • Multimodal Adapter for Improved Image Understanding and Generation
    • Connected Qwen2.5-VL to FLUX.1-schnell through a lightweight BLIP-2 Q-Former and trainable projection layers, bridging language and diffusion representations.
  • Query-Aware KV Cache Compression for Efficient Inference
    • Extended Hugging Face DynamicCache to assign separate KV retention budgets to text, image, and video tokens using attention scores, query similarity, and recency.

Fun Projects

 

images/publication_preview/smilecoin.png
SmilecoinπŸ˜ƒ: A minimal Bitcoin-style blockchain with proof-of-work mining, ASCII visualization, and LAN networking.

 

images/publication_preview/hri.png
A browser object delivery game for studying human robot interaction for RL (hierarchical POMDP) with Theory of Mind.

 

images/publication_preview/abm.png
An agent-based-model for studying effects of different precautions in Covid-19.

Experience

  • Sep. 2026 – Present
    Research Scientist Intern
    Adobe
    San Jose, CA
    • Enabled personalized, photorealistic style transfer for Adobe Lightroom users from only a few reference images.
    • Developed and trained a style encoder that distills the images into compact style tokens.
  • Jun. 2026 – Aug. 2026
    Research Assistant
    NEC
    Princeton, NJ (Remote)
    • Designed a compact scene representation that distills multimodal (RGB + LiDAR) Gaussian Splatting into a single differentiable signed distance field that returns geometry, semantics, physical properties, and calibrated uncertainty from a single query.
    • Validated the field distillation on the FOCI benchmark.
  • May 2026 – Aug. 2026
    Research Scientist Intern
    NVIDIA
    Santa Clara, CA
    • Developed Clean Forcing, an efficient method for reducing drift and improving long-horizon quality in autoregressive video and world models; filed a patent and integrated it into FlashDreams.
    • Enhanced the interactability of OmniDreams, a Cosmos Predict 2.5-based autonomous driving world model, with prompt hot-swapping, restyle LoRAs, Clean Forcing, and conditioned object insertion.
  • Sep. 2025 – Dec. 2025
    Research Scientist Intern
    Samsung
    Mountain View, CA
    • Improved spatial and temporal consistency in camera- and motion-controllable video generation built on Diffusion Forcing Transformers (DFoT) with history guidance.
    • Designed and evaluated new architectures for DFoT-based video diffusion models.
  • Jun. 2025 – Aug. 2025
    Research Scientist Intern
    Socure
    Boston, MA
    • Post-trained the Qwen-Image-Edit model to generate synthetic ID images, including text and profile photos, to enhance ID fraud detection.
    • Trained an ID fraud detection model using synthetic ID images and improved its robustness with a StyleGAN3-based evaluator.
  • May 2024 – Aug. 2024
    Research Scientist Intern
    bitHuman
    Boston, MA
    • Implemented an audio-driven 3D emotional talking-head generation framework.
    • Applied a variational autoencoder to generate dynamic facial landmarks driven by input audio.
  • Sep. 2023 – May 2024
    Research Assistant
    Northeastern University, Synergetic Media Learning Lab
    Boston, MA
    • Developed a text-to-3D Gaussian Splatting pipeline to simulate physics-grounded motion across diverse materials and object properties.
    • Integrated an LLM with evolutionary search and chain-of-thought (CoT) to iteratively optimize prompts, generating more precise 3D objects.
  • Sep. 2022 – Sep. 2023
    Research Assistant
    Northeastern University, Cognitive Embodied Social Agents Research Lab
    Boston, MA
    • Formulated hierarchical deep Bayesian reinforcement learning methods for Partially Observable Markov Decision Processes (POMDPs).
    • Integrated Theory of Mind into the POMDP framework to enhance decision-making processes.
  • Jun. 2021 – Jun. 2022
    Research Assistant
    University of California, Davis, FoxLab
    Davis, CA
    • Applied a Mask R-CNN model to detect 2D keypoints for tracking primates (Rhesus macaques).
    • Enabled the model to recognize primate behaviors through analysis of gestures and body features.

Education


Grants

Modal Neurips $10,000 GPU Grant, 2026

Modal ICLR $10,000 GPU Grant, 2026

ICCV Travel Grant, 2025

FG PhD Doctoral Consortium Travel Grant, 2025


Teaching

Courses TA'd: CS5100 Foundations of Artificial Intelligence; CS5150/4150 Game Artificial Intelligence; CS5340: Computer/Human Interaction; CS3200: Intro to Database Systems


Academic Services

Conference Reviewer: CVPR, NeurIPS, ICLR, ICML, ICCV, ECCV, AAAI, FG, MM

Journal Reviewer: Neural Computing and Applications, TPAMI, TKDD, Neural Networks, IEEE TNNLS, Knowledge‑Based Systems, AI Reviews, Multimedia Tools and Applications