Executive Summary

This research initiative investigates how embodied artificial intelligence can be redesigned to support low‑carbon, energy‑efficient autonomous mobility. The work responds to a growing gap in current AV development: while safety and reliability have advanced, the computational and operational sustainability of AI‑driven mobility remains largely unaddressed.

The study integrates reinforcement learning, sustainable computing, and transportation systems analysis to develop an autonomous driving agent that learns eco‑driving behaviours while reducing the energy footprint of AI training and inference. By combining compute‑efficient model architectures with energy‑aware decision‑making, the project aims to demonstrate measurable reductions in vehicle‑level energy use and system‑level emissions.

The expected impact spans technical innovation, environmental performance, and policy relevance—positioning embodied AI as a strategic enabler of climate‑aligned mobility in the UK and EU.

Focus and Approach

Core Research Domains

  • Compute‑efficient reinforcement learning for embodied AI

  • Energy‑aware sensorimotor integration for autonomous driving

  • Eco‑driving behaviour modelling and system‑level sustainability analysis

  • Climate‑resilient operational design for AVs

Methodological Philosophy The study adopts a simulation‑first, mixed‑methods research design combining experimental RL development, quantitative performance evaluation, and policy analysis. The approach emphasises iterative prototyping, measurable efficiency gains, and alignment with real‑world regulatory frameworks.

Key Findings (Interim)

  • Structural Challenge:
  • Current AV learning pipelines rely on large, energy‑intensive neural networks, creating a misalignment between technological progress and sustainability goals.

  • Opportunity:
  • Curiosity‑driven exploration significantly reduces data requirements, enabling more efficient learning in complex or sparse‑reward driving environments.

  • Design Implication:
  • Integrating real‑time energy consumption into the reward structure shifts driving behaviour toward smoother acceleration, anticipatory braking, and congestion‑aware routing.

  • System Insight:
  • Microscopic eco‑driving behaviours, when deployed across multiple agents, show early indications of reducing network‑wide delay and emissions in simulation.

  • Policy Gap:
  • Existing AV regulatory frameworks lack explicit sustainability metrics, leaving a critical opportunity for compute‑efficiency and carbon‑reporting standards.

    Policy Recommendations (Preliminary)

  • Establish Sustainability Reporting Requirements:
  • Mandate energy‑per‑inference and training‑carbon‑intensity disclosures for AV systems entering regulatory sandboxes.

  • Integrate Eco‑Driving Metrics into Type Approval:
  • Include energy‑per‑kilometre and congestion‑impact indicators in AV performance assessments.

  • Adopt Compute‑Efficiency Benchmarks:
  • Encourage the use of quantized, sparse, or otherwise efficient model architectures in safety‑critical AI applications.

  • Support Climate‑Resilient Testing Protocols:
  • Expand ODD definitions to include adverse‑weather scenarios and require performance validation under these conditions.

    1. Introduction

    Autonomous mobility is emerging as a transformative force in transportation, yet its environmental implications remain underexamined. While AVs promise efficiency and safety, the energy demands of modern AI systems and the operational behaviours of autonomous fleets could unintentionally increase emissions if not carefully managed.

    This research addresses the need for AI systems that are not only intelligent and safe but also aligned with national decarbonisation strategies. It explores how embodied AI can be redesigned to minimise energy use across both computation and vehicle operation.

    1.1 Context for This Work

    • Rising transportation emissions and pressure to meet net‑zero targets

    • Increasing computational costs associated with deep learning models

    • Growing interest in eco‑driving and congestion‑reduction strategies

    • Regulatory momentum toward sustainable and climate‑resilient mobility systems

    1.2 Report Structure

    • Overview of the research problem and context

    • Description of the methodological approach

    • Summary of qualitative and quantitative insights

    • Preliminary policy and design recommendations

    2. Approach to the Study

    The research follows a structured, four‑phase plan:

    1. Architecture Development: Designing a compute‑efficient reinforcement learning agent using quantized neural networks and curiosity‑driven exploration.

    2. Energy‑Aware Training: Integrating real‑time energy estimation into the reward function to shape sustainable driving behaviours.

    3. System‑Level Evaluation: Deploying trained agents in multi‑agent traffic simulations to assess congestion, emissions, and energy impacts.

    4. Policy Analysis: Mapping findings to UK/EU regulatory frameworks and identifying opportunities for sustainability‑aligned governance.

    2.1 Approach to Qualitative Research

    • Literature Synthesis: Identifies gaps in embodied AI, sustainable computing, and transport decarbonisation strategies.

    • Policy Review: Examines UK and EU mobility frameworks to understand regulatory readiness for sustainability‑driven AI.

    • Expert Consultation (planned): Provides insight into operational constraints, safety requirements, and industry adoption barriers.

    These methods help uncover structural, regulatory, and behavioural factors shaping the feasibility of sustainable AV systems.

    2.2 Approach to Quantitative Analysis and Key Findings

    Quantitative Methods

    • Reinforcement learning experiments in CARLA

    • Energy‑use estimation using vehicle dynamics models

    • Compute‑efficiency measurement (FLOPs, energy per training step)

    • Multi‑agent traffic simulation using SUMO

     

    Early Measurable Insights

    • Quantized RL models show promising reductions in compute cost during inference.

    • Energy‑aware reward shaping leads to smoother driving profiles and lower per‑kilometre energy use.

    • Multi‑agent deployment indicates potential reductions in network‑wide delay and emissions intensity.