Efficient Vision Language Action (VLA) Models for Autonomous Driving
05.10.2026, Diplomarbeiten, Bachelor- und Masterarbeiten
MS Thesis on VLA for autonomous driving in collaboration with DENSO Automotive
Vision Language Action (VLA) models show strong performance for data-driven autonomous driving. As VLA systems combine LLMs for text input, computer vision encoder and complex action generation, the system often consume considerable resources and memory. A recent line of research is focusing on more efficient VLA models, mainly by optimizing how the different modules VLAs interact.
This master’s thesis focuses on evaluation, benchmarking and enhancing of VLA models which focus on efficiency The project involves systematic evaluation of different VLA systems and optimization techniques on realistic benchmarks.
Key Research Areas
- Review of VLA models for autonomous driving regarding efficiency
- Quantitative evaluation of VLA systems and features regarding inference efficiency (time, memory)
- (optional) Enhancements of VLA systems for inference, integrating or combining different mechanisms for efficiency improvement
Technical Prerequisites (or Motivation to Learn)
- Interest in autonomous driving, machine learning
- Basic understanding of deep learning and model training
- Ability to read and modify research codebases
- Proficiency in Python
- Experience in VLA systems is a plus
Kontakt: christian.prehofer@tum.de


