INTERNSHIP DETAILS

Master thesis - LiDAR-Enhanced Vision-Language-Action Models for Safer Autonomous Truck

CompanyScania
LocationSödertälje kommun
Work ModeOn Site
PostedOctober 2, 2026
Internship Information
Core Responsibilities
The student will investigate the integration of LiDAR data into vision-language-action models to improve spatial reasoning and safety for autonomous trucks. This involves developing fusion mechanisms, fine-tuning models, and evaluating performance across various driving scenarios.
Internship Type
full time
Company Size
326
Visa Sponsorship
No
Language
English
Working Hours
40 hours
Apply Now →

You'll be redirected to
the company's application page

About The Company
Scania is een toonaangevende producent van bedrijfsauto's, bussen en industrie- en scheepsmotoren. De fabriek in Zwolle is de grootste Scaniafabriek ter wereld. Voor meer informatie www.scania.nl
About the Role

30 credits - LiDAR-Enhanced Vision-Language-Action Models for Safer Autonomous Trucks 


Introduction 

The VINNOVA-FFI Multi-modal Open-set Perception for Safer Autonomous Trucks (MOPS) project investigates how multimodal sensing and open-set perception can improve the safety of autonomous heavy-duty vehicles. Autonomous trucks must operate in open-world environments containing rare, unfamiliar, or previously unseen objects and situations. Their size, braking distance, limited maneuverability, and operational environments place particularly high demands on reliable perception and spatial reasoning. 


Background 

Camera-based end-to-end driving models capture rich semantic information, but their estimates of geometry, depth, and object distance may be sensitive to poor visibility, occlusion, and visually ambiguous scenes. LiDAR provides complementary three-dimensional measurements that could improve geometric understanding and help a driving model respond safely to unfamiliar objects and scenarios. 

NVIDIA Alpamayo is a family of open vision-language-action models for autonomous driving. Its models process multi-camera observations and driving context to generate trajectories together with interpretable Chain-of-Causation reasoning traces. Current open releases primarily rely on camera video and ego-motion inputs, creating an opportunity to investigate how LiDAR can be incorporated into the model’s perception, reasoning, and action-prediction pipeline. 

The thesis will initially rely on open-source Alpamayo models and training recipes and suitable datasets containing synchronized camera, LiDAR, and vehicle-motion data, such as the NVIDIA Physical AI Autonomous Vehicles dataset. 



Objective 

The thesis will investigate how LiDAR information can be integrated into Alpamayo to improve spatial reasoning, trajectory prediction, and robustness to unfamiliar objects and situations. 


Possible research directions include: 

  • representing LiDAR data as point-cloud, voxel, range-image, bird’s-eye-view, or learned tokens suitable for a vision-language-action model; 

  • comparing early, intermediate, and late fusion of camera and LiDAR features; 

  • developing token-efficient fusion mechanisms that limit the additional memory and inference cost; 

  • fine-tuning a suitable open Alpamayo release using synchronized camera and LiDAR observations; 

  • evaluating whether LiDAR improves trajectory prediction and spatial reasoning around previously unseen objects; 

  • studying whether the model’s reasoning traces correctly identify uncertainty, unfamiliar objects, and safety-relevant geometric constraints; 

  • evaluating robustness under poor visibility, occlusion, sensor degradation, or missing modalities; or 

  • analyzing which truck-relevant driving scenarios benefit most from explicit three-dimensional information. 


The proposed multimodal model should be compared with a camera-only Alpamayo baseline. Evaluation may include trajectory accuracy, collision-related metrics, open-set detection or response metrics, reasoning–action consistency, robustness, and computational efficiency. Where suitable data and benchmarks are available, the study may examine scenarios particularly relevant to heavy-duty vehicles, such as long stopping distances, narrow clearances, long-range perception, unusual road users, and unfamiliar obstacles. 

The exact model variant, fusion strategy, dataset, and evaluation framework will be finalized together with the selected student and the MOPS project team, based on available open-source tools, computational resources, and the project’s priorities. 



The Project Offers  

The student will contribute to the MOPS industrial research project and work on a problem combining multimodal machine learning, autonomous driving, three-dimensional perception, open-set recognition, and vision-language-action modeling. The work is expected to produce a reusable LiDAR-enhanced Alpamayo baseline, sensor-fusion component, or experimental framework that can support continued MOPS research. 

The student will receive supervision from TRATON and have opportunities to interact with the project’s academic and industrial collaborators. 



Who are we looking for? 

We are looking for a Master’s student in computer science, machine learning, robotics, engineering physics, electrical engineering, or a related field. 


A suitable candidate should have: 

  • strong programming skills, preferably in Python and PyTorch; 

  • knowledge of machine learning and deep learning; 

  • an interest in autonomous driving, multimodal transformers, computer vision, or three-dimensional perception; 

  • experience with LiDAR, point-cloud processing, open-set recognition, or large-scale model fine-tuning is beneficial but not required; and 

  • motivation to combine scientific investigation with practical implementation. 


The planned thesis start is January 2027. 

Number of students: 1 

Start date for the thesis work: [To be agreed] 

Estimated time required: 20 weeks, full time (30 credits) 



Contact persons and supervisors 

Thomas Gustafsson, thomas.gustafsson@scania.com; Jesper Eriksson, jesper.ericsson@scania.com;  

Hiring Manager: Maria Linnarsson, maria.linnarsson@scania.com 

 


Application 

Your application must include a CV, personal letter, and transcript of grades. 

A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified.

 

Publication date:
1.10.2026 - 30.11.2026 (applications evaluated continuously)
Key Skills
PythonPyTorchMachine LearningDeep LearningAutonomous DrivingMultimodal TransformersComputer VisionThree-dimensional PerceptionLiDARPoint-cloud ProcessingOpen-set RecognitionLarge-scale Model Fine-tuningSpatial ReasoningTrajectory PredictionSensor Fusion
Categories
TechnologyEngineeringScience & ResearchSoftwareTransportation