Master thesis - Evaluating and Improving Explainable Vision-Language Models for Autonomous driving

You'll be redirected to
the company's application page
30 credits - Evaluating and Improving Explainable Vision-Language Models for Autonomous Driving
Introduction
A Master's thesis is an excellent way to get closer to TRATON Group R&D and build relationships for the future.
Background
Vision-language models (VLMs) are increasingly used in autonomous-driving research to interpret traffic scenes, answer questions, describe relevant objects, and provide natural-language explanations for driving decisions. These explanations could support model development, debugging, safety analysis, and communication between automated vehicles and human operators.
However, a plausible explanation is not necessarily a faithful explanation. A VLM may produce fluent reasoning that is not grounded in the visual observations, omit safety-critical objects, hallucinate scene elements, or provide an explanation that is inconsistent with its predicted decision. Conventional language-generation metrics may reward similarity to a reference answer without revealing whether the explanation reflects the model’s actual reasoning.
The thesis will initially rely on public autonomous-driving datasets, open-source VLMs, and language-based driving benchmarks such as DriveLM and LingoQA. Depending on the selected direction, the work may investigate models that explain scene understanding, behavior prediction, trajectory planning, or end-to-end driving decisions.
Objective
The thesis will investigate how the explainability of vision-language models for autonomous driving can be evaluated and improved. The work should distinguish between explanations that are merely plausible and explanations that are grounded in the observed scene and faithful to the model’s output.
Possible research directions include:
-
defining measurable properties of useful driving explanations, such as correctness, visual grounding, faithfulness, completeness, uncertainty awareness, and decision consistency;
-
evaluating whether generated explanations refer to the correct objects, road users, traffic rules, and spatial relationships;
-
comparing free-form natural-language explanations with structured explanations based on scene graphs, object references, concepts, or causal relations;
-
grounding explanations in image regions, object tracks, bird’s-eye-view representations, or other intermediate model features;
-
comparing intrinsic explanation mechanisms with post-hoc methods applied to existing VLMs;
-
using visual perturbations or counterfactual scenes to test whether explanations respond to safety-relevant changes;
-
detecting hallucinated, unsupported, or internally inconsistent explanations;
-
developing training objectives that improve agreement between visual evidence, predicted behavior, and generated explanations;
-
investigating calibrated explanations that communicate uncertainty or abstain when the available evidence is insufficient; or
-
studying whether explanations help engineers or human operators identify model failures more efficiently.
The selected approach should be evaluated on at least one open VLM and autonomous-driving dataset. Evaluation may include visual question-answering performance, object and relation grounding, hallucination rate, explanation–decision consistency, sensitivity to controlled scene changes, uncertainty calibration, and computational cost. Human evaluation may also be used to assess whether the explanations are understandable and useful for diagnosing model behavior.
A central challenge will be establishing whether an explanation reflects the information that influenced the model, rather than being a convincing description generated after the decision. The thesis should therefore include evaluation methods that go beyond text similarity and test the relationship between visual evidence, model output, and explanation.
The exact model, driving task, dataset, and explanation method will be finalized together with the selected student, based on the student’s background, available open-source tools, computational resources, and the expected scientific contribution.
The Project Offers
The student will work on a research problem combining machine learning, autonomous driving, multimodal models, and explainable artificial intelligence. The work is expected to produce a reusable evaluation framework, explainability method, or model baseline for studying the reliability of VLM-generated driving explanations.
The student will receive supervision from TRATON and have opportunities to interact with researchers and engineers working on machine learning and autonomous driving.
Who are we looking for?
-
We are looking for a Master’s student in computer science, machine learning, robotics, engineering physics, electrical engineering, or a related field.
-
A suitable candidate should have:
-
strong programming skills, preferably in Python and PyTorch;
-
knowledge of machine learning and deep learning;
-
an interest in autonomous driving, computer vision, transformers, multimodal models, or explainable AI;
-
experience with vision-language models, natural-language processing, model interpretability, or autonomous-driving datasets is beneficial but not required; and
-
motivation to combine scientific investigation with practical implementation.
The planned thesis start is January 2027.
Number of students: 1
Start date for the thesis work: [To be agreed]
Estimated time required: 20 weeks, full time (30 credits)
Contact persons and supervisors
Thomas Gustafsson, thomas.gustafsson@scania.com;
Hiring Manager: Maria Linnarsson, maria.linnarsson@scania.com
Application
Your application must include a CV, personal letter, and transcript of grades.
A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified.
Prep Tools
PROFESSIONAL COVER LETTER TEMPLATES
Template Library
Professional templates
50+ templates for every role
STUCK ON A QUESTION? PRACTICE IT
Practice Any Question
Get instant AI feedback
"How would you design a scalable system for Scania's use case?"
ACE YOUR INTERVIEW IN REAL-TIME
Silent AI Co-Pilot
Real-time interview help
"Why Scania?"
💡 Mention their Automotive and your passion for Python