Haodong Zhang 张浩栋

My research centers on learning coherent world representations from multimodal data with limited supervision. I study how observations of the same objects and events can be connected across modalities while preserving their complementary information. Building on these representations, I aim to develop hierarchical world models that link local dynamics and control to broader planning and coordination. My long-term goal is to enable modular robots to act independently or assemble for different tasks, with learned models that can be reused and composed as their physical structure changes. I am a Research Assistant in Prof. Randall Balestriero's galilAI lab at Brown University.

Self-Supervised Learning
Learning useful representations from unlabeled data.
Multimodal Learning
Aligning modalities while preserving complementary information.
Continued Pretraining
When further pretraining improves transfer.
Robotics
Shared world models for cooperative robot planning.
Haodong Zhang

Education

  1. M.S. in Computer Science

    Brown University

  2. B.S. in Computer Science

    University of California, Irvine

News

  1. 🎉 Graduated from Brown University with an M.S. in Computer Science.

  2. The fin-whale vocalization detection paper was listed as an oral at the NeurIPS 2025 Workshop on AI for Non-Human Animal Communication.

  3. 🎉 Graduated from UC Irvine with a B.S. in Computer Science, Magna Cum Laude†.

    † GPA ranked in the top 5%.

  4. The weakly supervised self-ensembling Vision Transformer paper appeared at the 2023 IEEE Conference on Artificial Intelligence.

Selected Publications

Initial feature geometry and continued-pretraining response across four pretrained encoders.Submitted 2026

When Does Continued Pretraining Improve Frozen Visual Representations?

Haodong Zhang, J. Ambsdorf

Studies when continued pretraining improves frozen visual features and how representation geometry predicts gains across objectives, encoders, data budgets, and datasets.

Conceptual schematic of two modality distributions mapped into shared semantic neighborhoods while retaining complementary structure.In Progress

How Much to Align in Multimodal Learning? Unleashing Synergy without Contrastive Alignment

Haodong Zhang, H. Wu

A non-contrastive multimodal objective that coordinates shared structure while preserving modality-specific information, allowing fused representations to use complementary evidence.

Retrieval performance of three alignment heads across seven frozen encoders.In Progress

From Encoded to Accessible: Pretraining Wires and Alignment Routes in Frozen Vision Representations

Haodong Zhang, Z. Liu, Z. Li

Separates information encoded in frozen vision backbones from what alignment and downstream tasks can access, revealing channels reactivated by targeted supervision.

Six self-supervised methods compared with supervised learning across 41 image datasets.Under Review 2026

Is SSL Ready for In-Domain Pretraining? A Cross-Dataset Benchmark and Analysis

S. BuGhanem, L. Hu, Haodong Zhang, , R. Balestriero

Benchmarks six self-supervised methods across 41 image datasets to study when limited-data in-domain pretraining competes with supervised learning.

Figure 3: SSL embeddings of fin-whale pulses and non-pulses across four signal-to-noise thresholds in the Seglvik dataset.NeurIPS Workshop 2025 (Oral)

Supervised vs Self-Supervised Representation Learning for Fin Whale Vocalization Detection

A. Chareyre*, Haodong Zhang*, S. Ge, R. Balestriero, H. Glotin, S. Paris

* Equal contribution.

Learns fin-whale acoustic representations from unlabeled recordings for robust detection with few labels, low signal-to-noise ratios, and shifts across recording sites.

Teacher-student U-shaped Transformers with scribble supervision, transformation consistency, and exponential moving-average updates for cardiac MRI segmentation.IEEE CAI 2023

Weakly-Supervised Self-Ensembling Vision Transformer for MRI Cardiac Segmentation

Z. Wang, Haodong Zhang, Y. Liu

Combines scribble supervision, self-ensembling, and transformation consistency in a Vision Transformer for cardiac MRI segmentation with sparse annotations.

Selected Projects

Sequential world model architecture for cooperative robot manipulation.

DINO-SeqWM: Multi-Robot Cooperation with Shared and Sequential World Model

2026-01 – 2026-05

Sequential world models over frozen DINOv3 features let two robot arms plan in a shared latent future for cooperative manipulation.

Food segmentation, depth estimation, and nutrition estimation pipeline.

3D-FoodCalorie: Geometry-Aware Calorie Estimation from Single Image

2025-01 – 2025-05

Combines food segmentation, monocular depth estimation, and RGB-D fusion to estimate food calories from a single image.

Image denoising results using a pretrained diffusion model.

Pre-trained Diffusion Model for Image Denoising with Contrastive Learning

2024-09 – 2024-12

Uses a contrastive timestep estimator to place a noisy image within a pretrained diffusion process, then reverses that process for denoising.

Get in touch

Let’s explore what comes next.

haodong_zhang@brown.edu

I welcome conversations about representation learning, world models, and ideas for collaboration.