INTERNSHIP DETAILS

Master thesis - Evaluating and Improving Explainable Vision-Language Models for Autonomous driving

CompanyScania
LocationSödertälje kommun
Work ModeOn Site
PostedOctober 2, 2026
Internship Information
Core Responsibilities
The student will investigate how to evaluate and improve the explainability of vision-language models within the context of autonomous driving. This involves developing a reusable evaluation framework or model baseline that distinguishes between plausible and faithful explanations.
Internship Type
full time
Company Size
326
Visa Sponsorship
No
Language
English
Working Hours
40 hours
Apply Now →

You'll be redirected to
the company's application page

About The Company
Scania is een toonaangevende producent van bedrijfsauto's, bussen en industrie- en scheepsmotoren. De fabriek in Zwolle is de grootste Scaniafabriek ter wereld. Voor meer informatie www.scania.nl
About the Role

30 credits - Evaluating and Improving Explainable Vision-Language Models for Autonomous Driving 

 

Introduction
 

A Master's thesis is an excellent way to get closer to TRATON Group R&D and build relationships for the future. 



Background 

Vision-language models (VLMs) are increasingly used in autonomous-driving research to interpret traffic scenes, answer questions, describe relevant objects, and provide natural-language explanations for driving decisions. These explanations could support model development, debugging, safety analysis, and communication between automated vehicles and human operators. 

However, a plausible explanation is not necessarily a faithful explanation. A VLM may produce fluent reasoning that is not grounded in the visual observations, omit safety-critical objects, hallucinate scene elements, or provide an explanation that is inconsistent with its predicted decision. Conventional language-generation metrics may reward similarity to a reference answer without revealing whether the explanation reflects the model’s actual reasoning. 

The thesis will initially rely on public autonomous-driving datasets, open-source VLMs, and language-based driving benchmarks such as DriveLM and LingoQA. Depending on the selected direction, the work may investigate models that explain scene understanding, behavior prediction, trajectory planning, or end-to-end driving decisions. 



Objective 

The thesis will investigate how the explainability of vision-language models for autonomous driving can be evaluated and improved. The work should distinguish between explanations that are merely plausible and explanations that are grounded in the observed scene and faithful to the model’s output. 

Possible research directions include: 

  • defining measurable properties of useful driving explanations, such as correctness, visual grounding, faithfulness, completeness, uncertainty awareness, and decision consistency; 

  • evaluating whether generated explanations refer to the correct objects, road users, traffic rules, and spatial relationships; 

  • comparing free-form natural-language explanations with structured explanations based on scene graphs, object references, concepts, or causal relations; 

  • grounding explanations in image regions, object tracks, bird’s-eye-view representations, or other intermediate model features; 

  • comparing intrinsic explanation mechanisms with post-hoc methods applied to existing VLMs; 

  • using visual perturbations or counterfactual scenes to test whether explanations respond to safety-relevant changes; 

  • detecting hallucinated, unsupported, or internally inconsistent explanations; 

  • developing training objectives that improve agreement between visual evidence, predicted behavior, and generated explanations; 

  • investigating calibrated explanations that communicate uncertainty or abstain when the available evidence is insufficient; or 

  • studying whether explanations help engineers or human operators identify model failures more efficiently. 


The selected approach should be evaluated on at least one open VLM and autonomous-driving dataset. Evaluation may include visual question-answering performance, object and relation grounding, hallucination rate, explanation–decision consistency, sensitivity to controlled scene changes, uncertainty calibration, and computational cost. Human evaluation may also be used to assess whether the explanations are understandable and useful for diagnosing model behavior. 

A central challenge will be establishing whether an explanation reflects the information that influenced the model, rather than being a convincing description generated after the decision. The thesis should therefore include evaluation methods that go beyond text similarity and test the relationship between visual evidence, model output, and explanation. 

The exact model, driving task, dataset, and explanation method will be finalized together with the selected student, based on the student’s background, available open-source tools, computational resources, and the expected scientific contribution. 



The Project Offers 

The student will work on a research problem combining machine learning, autonomous driving, multimodal models, and explainable artificial intelligence. The work is expected to produce a reusable evaluation framework, explainability method, or model baseline for studying the reliability of VLM-generated driving explanations. 

The student will receive supervision from TRATON and have opportunities to interact with researchers and engineers working on machine learning and autonomous driving. 



Who are we looking for? 

  • We are looking for a Master’s student in computer science, machine learning, robotics, engineering physics, electrical engineering, or a related field. 

  • A suitable candidate should have: 

  • strong programming skills, preferably in Python and PyTorch; 

  • knowledge of machine learning and deep learning; 

  • an interest in autonomous driving, computer vision, transformers, multimodal models, or explainable AI; 

  • experience with vision-language models, natural-language processing, model interpretability, or autonomous-driving datasets is beneficial but not required; and 

  • motivation to combine scientific investigation with practical implementation. 


The planned thesis start is January 2027. 

 

Number of students: 1 

Start date for the thesis work: [To be agreed] 

Estimated time required: 20 weeks, full time (30 credits) 

Contact persons and supervisors 

Thomas Gustafsson, thomas.gustafsson@scania.com; 

Hiring Manager: Maria Linnarsson, maria.linnarsson@scania.com  



Application 

Your application must include a CV, personal letter, and transcript of grades. 

A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified. 


Publication date:
1.10.2026 - 30.11.2026 (applications evaluated continuously)
Key Skills
PythonPyTorchMachine LearningDeep LearningAutonomous DrivingComputer VisionTransformersMultimodal ModelsExplainable AINatural Language ProcessingModel InterpretabilityData AnalysisResearchProgramming
Categories
Science & ResearchTechnologyEngineeringSoftwareData & Analytics