Direkt zum Inhalt springen
login.png Login    |
de | en
MyTUM-Portal
Technische Universität München

Technische Universität München

Sitemap > Schwarzes Brett > Abschlussarbeiten, Bachelor- und Masterarbeiten > Master Thesis: 4D JEPA (Joint Embedding Predictive Architectures) for Cardiac MRI
auf   Zurück zu  Nachrichten-Bereich      Browse in News  nächster    

Master Thesis: 4D JEPA (Joint Embedding Predictive Architectures) for Cardiac MRI

09.10.2026, Abschlussarbeiten, Bachelor- und Masterarbeiten

Beyond next-token prediction: Predicting embeddings

Chair: Chair for Computational Imaging and AI in Medicine (CompAI), TU Munich
Professor: Prof. Dr. Julia Schnabel
Advisors: Marta Hasny, Sameer Ambekar
Type: Master Thesis
Contact: marta.hasny@tum.de, sameer.ambekar@tum.de

Introduction

Joint Embedding Predictive Architectures (JEPA) [1,2,3] offer a compelling alternative to pixel-reconstruction pretraining. Standard Transformer-based pretraining objectives, such as autoregressive next-token prediction in language models and pixel-level reconstruction in vision, are designed to recover the input signal. JEPA instead predicts the representation of a masked region in latent space. This makes the model invariant to high-frequency detail that carries no semantic content. Its video extension, V-JEPA [2], learns motion-sensitive representations by predicting across time, and so generalises to spatio-temporal data.

Cine Cardiac Magnetic Resonance (CMR) [4] imaging is intrinsically four-dimensional: a three-dimensional volume of the heart captured across the cardiac cycle. Yet the field has not converged on how to present this structure to deep learning models. Simpler formats, such as 2D slices or key frames, are economical but discard motion or spatial context. Richer formats, such as full 3D or 4D volumes, preserve this information at considerable computational cost.

JEPA is well suited to this trade-off. Rather than reconstructing pixel-level detail, it predicts how a representation evolves across space and time. This matches clinical practice, where diagnosis depends on tracking wall motion and chamber volume rather than on intensities corrupted by scanner noise. No prior work applies this paradigm to cine CMR in its native four-dimensional form. Existing approaches rely on key frames alone, or embed the full sequence into architectures designed for 3D input, which leaves open how much clinically relevant information is discarded.

Reconstruction-based pretraining [5,6] compounds this problem, as decoding full 4D volumes is where its computational cost is greatest. By predicting in latent space, JEPA circumvents this cost, and this thesis develops the first pretraining scheme to operate natively on full 4D cardiac data. The benefit of pretraining for cardiac imaging is by now well established. What is still missing is a principled account of how the data's own structure should be used in the model.

Objectives

  • The first comprehensive comparison of full 4D input against end-diastolic and end-systolic (ED/ES) key frames, with architecture and data held fixed, isolating dimensionality as the sole determinant of any performance gap.
  • A standardized evaluation of 2D, 2D+T, 3D, and 4D CMR formats on segmentation and functional tasks under matched architectures, establishing which format justifies its computational cost.
  • Adapting Joint Embedding Predictive Architectures (JEPA) to 4D cine CMR to learn latently features without pixel-level reconstruction, so that model capacity targets cardiac motion rather than scanner noise.

Requirements

  • Good understanding of deep learning and, ideally, vision-language models.
  • Practical experience with Python and PyTorch.
  • Interest in medical imaging research; familiarity with representation learning is an advantage.

Application

Submit your application at the TUM Thesis Management Portal.

References

[1] LeCun, Y. A Path Towards Autonomous Machine Intelligence, 2022.
[2] Bardes, A. et al. Revisiting Feature Prediction for Learning Visual Representations from Video (V-JEPA), 2024.
[3] Assran, M. et al. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (I-JEPA), 2023.
[4] Chen, C., et al. Deep Learning for Cardiac Image Segmentation: A Review, 2020.
[5] He, K. et al. Masked Autoencoders Are Scalable Vision Learners, 2022.
[6] Tong, Z. et al. VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training, 2022.

Kontakt: sameer.ambekar@tum.de

Mehr Information

https://thesis.aet.cit.tum.de/topics/df10a736-1119-41e2-b4a3-41f75b68a217

Termine heute

no events today.

Veranstaltungskalender