The OBD-AI project proposes an agentic AI diagnostic platform that combines vehicle telemetry, diagnostic trouble codes (DTCs), service information, historical repair knowledge, machine learning and large language models.

The central research question is:

How can an AI system combine authoritative automotive knowledge with real-time vehicle data while remaining grounded, explainable, auditable and safe?

This paper proposes that the answer should not be a conventional chatbot.

OBD-AI should instead be designed as a RAG-LLM + MCP agentic diagnostic system.

The architecture separates two fundamentally different capabilities:

  • RAG-LLM answers: What does the technical knowledge say?
  • MCP tools answer: What is the vehicle reporting right now, and what authorized system operation can be performed?
  • Agent orchestration answers: What evidence should be collected, what knowledge should be retrieved, and what diagnostic hypothesis best explains the evidence?

IAS RESEARCH | KEENSOFTWARE

Research & Development White Paper

Model Context Protocol (MCP) and Retrieval-Augmented Generation (RAG-LLM) for OBD-AI

An Agentic, Grounded and Safety-Governed Architecture for Intelligent Vehicle Diagnostics

Version: 1.0 — Research & Development Draft
Date: October 2026
Location: Winnipeg, Manitoba, Canada
Organizations: IAS-Research.com | KeenSoftware / KeenComputer.com | KeenDirect.com
Application Domain: ICE | Hybrid | PHEV | BEV | Fleet Diagnostics | Predictive Maintenance

Research status: This paper is an architectural and R&D proposal. Vehicle-protocol capabilities, OEM diagnostic access, service-manual licensing, MCP implementations, and production safety requirements must be validated against the exact hardware, software and standards versions used in deployment.

Executive Summary

The OBD-AI project proposes an agentic AI diagnostic platform that combines vehicle telemetry, diagnostic trouble codes (DTCs), service information, historical repair knowledge, machine learning and large language models.

The central research question is:

How can an AI system combine authoritative automotive knowledge with real-time vehicle data while remaining grounded, explainable, auditable and safe?

This paper proposes that the answer should not be a conventional chatbot.

OBD-AI should instead be designed as a RAG-LLM + MCP agentic diagnostic system.

The architecture separates two fundamentally different capabilities:

  • RAG-LLM answers: What does the technical knowledge say?
  • MCP tools answer: What is the vehicle reporting right now, and what authorized system operation can be performed?
  • Agent orchestration answers: What evidence should be collected, what knowledge should be retrieved, and what diagnostic hypothesis best explains the evidence?

The design therefore follows a fundamental principle:

RAG retrieves knowledge; MCP retrieves and operates on structured reality; the agent connects the two under explicit policy and safety controls.

For example, when a vehicle reports P0420, the AI should not simply ask an LLM to explain the code.

Instead:

  1. MCP reads the DTC.
  2. MCP retrieves VIN and vehicle configuration.
  3. MCP retrieves freeze-frame and relevant live PIDs.
  4. The diagnostic agent determines the diagnostic context.
  5. RAG retrieves the correct service-manual procedure.
  6. Hybrid/vector/graph retrieval identifies related components and failure modes.
  7. The LLM produces a diagnostic hypothesis.
  8. Evidence is attached to each recommendation.
  9. The system identifies the next verification test.
  10. Any vehicle-changing operation requires explicit authorization.

This transforms OBD-AI from an AI chatbot into an AI diagnostic engineering platform.

The research architecture also incorporates:

  • hybrid RAG;
  • metadata-filtered retrieval;
  • GraphRAG / Neo4j;
  • structured automotive knowledge graphs;
  • MCP tool servers;
  • agentic planning;
  • BDD/Cucumber regression testing;
  • human-in-the-loop verification;
  • provenance and citations;
  • cybersecurity controls;
  • telemetry privacy;
  • safety gates;
  • predictive maintenance;
  • fleet management;
  • digital-twin and simulation opportunities.

MCP is particularly relevant because the July 28, 2026 specification introduced a stateless protocol core, authorization hardening, caching of list results and other capabilities aimed at scalable agentic infrastructure. (Model Context Protocol Blog)

1. Research Objectives

The OBD-AI R&D program has six primary objectives.

Objective 1 — Ground AI diagnostic reasoning

Reduce hallucination by grounding responses in:

  • OEM service manuals;
  • technical service bulletins;
  • recall information;
  • diagnostic standards;
  • DTC definitions;
  • vehicle-specific procedures;
  • validated repair cases;
  • engineering knowledge.

RAG originated as an approach for combining parametric language-model knowledge with retrieved external knowledge, particularly for knowledge-intensive tasks. (UCL NLP)

Objective 2 — Connect AI to live vehicle information

The LLM cannot independently know:

  • current DTCs;
  • freeze-frame values;
  • vehicle VIN;
  • coolant temperature;
  • engine RPM;
  • fuel trims;
  • oxygen-sensor readings;
  • battery information;
  • trip history.

MCP provides a standardized mechanism for exposing external tools and resources to AI applications. The current MCP architecture includes tools, resources and prompts as core server primitives. (Model Context Protocol)

Objective 3 — Build an agentic diagnostic workflow

The system should not simply answer questions.

It should:

Observe → Retrieve → Correlate → Hypothesize → Verify → Recommend → Record

This represents a transition from:

Question answering

to:

AI-assisted diagnostic reasoning.

Objective 4 — Establish safety boundaries

The architecture must distinguish between:

Read-only operations

and

vehicle-changing operations.

For example:

Operation

Risk

Read VIN

Low

Read DTCs

Low

Read PIDs

Low

Read freeze frame

Low

Retrieve manual

Low

Create work order

Medium

Clear DTCs

High

Actuator test

High

ECU programming

Very high

HV-system control

Critical

This distinction must be implemented in software architecture—not merely described in an LLM prompt.

Objective 5 — Create measurable engineering quality

OBD-AI should be testable like conventional safety-sensitive software.

The project should therefore combine:

  • BDD;
  • Cucumber/Gherkin;
  • unit testing;
  • integration testing;
  • RAG evaluation;
  • MCP tool validation;
  • adversarial testing;
  • technician review;
  • regression testing.

Objective 6 — Establish a scalable commercial architecture

The platform should eventually support:

  • individual vehicle owners;
  • independent repair shops;
  • automotive technicians;
  • dealerships;
  • fleet operators;
  • used-vehicle inspection;
  • warranty organizations;
  • automotive engineering organizations;
  • EV service organizations.

2. The OBD-AI Research Hypothesis

The central hypothesis of this research is:

A vehicle diagnostic AI system will provide more reliable and actionable assistance when real-time structured vehicle evidence is obtained through controlled tools and domain knowledge is retrieved through an auditable RAG architecture, rather than relying on an LLM's parametric knowledge alone.

The architecture can therefore be represented as:

Diagnostic Intelligence=Vehicle Evidence+Domain Knowledge+Agentic Reasoning+VerificationDiagnostic\ Intelligence = Vehicle\ Evidence + Domain\ Knowledge + Agentic\ Reasoning + Verification

where:

Vehicle Evidence=MCP(OBD/CAN/ECU/Fleet)Vehicle\ Evidence = MCP(OBD/CAN/ECU/Fleet)

and:

Domain Knowledge=RAG(Manuals/TSBs/Repair Knowledge)Domain\ Knowledge = RAG(Manuals/TSBs/Repair\ Knowledge)

and:

Diagnostic Decision=LLM(Evidence+Retrieved Knowledge+Policy)Diagnostic\ Decision = LLM(Evidence + Retrieved\ Knowledge + Policy)

This is the foundation of OBD-AI.

3. Why Conventional OBD Diagnostics Are Insufficient

Traditional diagnostic workflows generally resemble:

Vehicle → Scan Tool → DTC → Technician

The scan tool is extremely useful, but a DTC is not necessarily a diagnosis.

For example:

P0420 — Catalyst System Efficiency Below Threshold

does not automatically mean:

Replace catalytic converter.

Possible causes can include:

  • exhaust leaks;
  • oxygen-sensor problems;
  • wiring;
  • air/fuel imbalance;
  • misfire history;
  • exhaust contamination;
  • catalyst degradation;
  • incorrect operating conditions;
  • other upstream faults.

Therefore:

DTC≠Root CauseDTC \neq Root\ Cause

A more realistic diagnostic relationship is:

DTC+VehicleContext+LiveEvidence+ServiceKnowledge→DiagnosticHypothesisDTC + VehicleContext + LiveEvidence + ServiceKnowledge \rightarrow DiagnosticHypothesis

This is precisely where RAG and MCP become important.

4. RAG-LLM: The Knowledge Layer

4.1 Basic architecture

The RAG architecture consists of:

Documents → Ingestion → Chunking → Embedding → Index → Retrieval → Reranking → LLM

For OBD-AI, this should become:

Manual → Document Intelligence → Automotive Metadata → Hybrid Retrieval → Evidence Set → Diagnostic Agent

5. OBD-AI Knowledge Corpus

The knowledge corpus should be divided into authority levels.

Tier 1 — Manufacturer information

  • service manuals;
  • workshop manuals;
  • diagnostic procedures;
  • wiring documentation;
  • technical service bulletins;
  • recalls;
  • manufacturer diagnostic information.

Tier 2 — Standards

Examples include:

  • SAE J1979;
  • SAE J2012;
  • ISO 15765;
  • ISO 14229;
  • ISO 15031;
  • related CAN/diagnostic standards.

SAE J1979 defines diagnostic test modes and their communication between OBD systems and test equipment; the current SAE listing identifies J1979_202505 as reaffirmed in May 2025. (SAE Mobilus)

Tier 3 — Engineering knowledge

  • textbooks;
  • training manuals;
  • engineering publications;
  • diagnostic theory;
  • component behavior.

Tier 4 — Curated field knowledge

  • validated repair cases;
  • technician observations;
  • internal troubleshooting notes;
  • fleet maintenance records.

Tier 5 — Community/public information

Potentially useful, but lower authority.

Such information should never automatically override an OEM procedure.

6. Automotive RAG Metadata

Metadata becomes one of the most important components of the system.

Each knowledge object should include:

Metadata

Example

Manufacturer

Toyota

Model

Prius

Model year

2018

Generation

Gen 4

Engine

2ZR-FXE

Powertrain

HEV

System

Emissions

ECU

Engine ECU

DTC

P0420

Component

Catalytic converter

Procedure

Catalyst efficiency test

Document

Service Manual

Section

Engine Control

Page

543

Source

OEM

Version

2025

Region

North America

License

Licensed

Confidence

Verified

The metadata filter should be applied before semantic ranking whenever possible.

This prevents a highly similar procedure for the wrong vehicle from winning the retrieval competition.

7. Hybrid RAG Architecture

OBD-AI should not rely on vector search alone.

The recommended architecture is:

Layer 1 — Lexical retrieval

Useful for:

  • P0420;
  • P0171;
  • sensor numbers;
  • connector IDs;
  • part numbers;
  • ECU names.

Layer 2 — Vector retrieval

Useful for:

  • natural-language questions;
  • symptom descriptions;
  • similar diagnostic procedures;
  • semantic relationships.

Layer 3 — Metadata filtering

Used for:

  • make;
  • model;
  • year;
  • engine;
  • powertrain;
  • system;
  • region.

Layer 4 — Reranking

A cross-encoder or equivalent model evaluates candidate passages.

Layer 5 — Graph retrieval

Neo4j/GraphRAG can connect:

Vehicle → ECU → DTC → Sensor → Component → Failure Mode → Procedure → Manual

GraphRAG research demonstrates the value of graph-based retrieval for connecting entities and answering questions across large private corpora. (arXiv)

8. Proposed OBD-AI Knowledge Graph

A representative ontology is:

Vehicle ├── hasVIN ├── hasEngine ├── hasPowertrain ├── containsECU │ └── DiagnosticSession ├── reportsDTC │ └── DTC │ ├── affectsSystem │ ├── relatesToComponent │ └── hasFailureMode │ ├── containsPID ├── containsFreezeFrame └── producesEvidence Component ├── hasSensor ├── hasFailureMode ├── hasInspectionProcedure └── referencedByManual

This allows questions such as:

Which components are associated with P0420 on this engine?

or:

Which previous cases showed this combination of P0171 + high fuel trim + MAF deviation?

Graph retrieval can complement conventional vector RAG rather than replacing it.

9. Model Context Protocol for OBD-AI

MCP provides the tool and context integration layer.

The current MCP documentation defines three important primitives:

  • Tools — executable functions;
  • Resources — contextual information;
  • Prompts — reusable interaction templates. (Model Context Protocol)

For OBD-AI, the most important primitive is the tool.

10. Proposed MCP Server Architecture

Instead of one enormous MCP server:

OBD-AI MCP └── everything

the recommended architecture is:

OBD-AI Agent | +----------+----------+ | | MCP Gateway RAG Gateway | | +------+------+ +-----+------+ | | | | | Telemetry Diagnostics Fleet Knowledge Graph MCP MCP MCP MCP DB | Vehicle / Puck

This gives each bounded context a clear security boundary.

11. Proposed MCP Servers

11.1 telemetry-mcp

Tools:

get_live_pids() get_freeze_frame() get_trip_history() get_sensor_history()

Purpose:

Provide real-time and historical vehicle measurements.

11.2 diagnostics-mcp

Tools:

read_dtcs() decode_dtc() get_readiness_monitors() get_diagnostic_status()

This server should primarily be read-only.

11.3 vehicle-info-mcp

Tools:

decode_vin() get_vehicle_profile() get_engine_configuration() get_powertrain_configuration() get_recall_status()

11.4 knowledge-mcp

Tools:

search_manuals() get_procedure() search_tsbs() get_source() get_citation()

This server becomes the controlled interface to RAG.

11.5 predictive-mcp

Tools:

get_component_health() forecast_failure() get_anomaly_score() get_maintenance_prediction()

11.6 fleet-mcp

Tools:

list_vehicles() get_service_history() get_vehicle_status() create_work_order() update_work_order()

Business-system writes should require their own authorization.

11.7 control-mcp

This is the most restricted server.

Potential tools:

clear_dtcs() run_actuator_test() execute_diagnostic_routine()

Future functionality might include ECU programming, but that should be treated as a substantially different safety and cybersecurity category.

12. RAG and MCP: Division of Responsibility

The most important architectural decision is:

Question

RAG

MCP

What does P0420 mean?

✓

✓ structured lookup

What is the manufacturer's procedure?

✓

MCP exposes RAG

What vehicle is connected?

 

✓

What DTCs are currently stored?

 

✓

What is coolant temperature now?

 

✓

What happened during last trip?

 

✓

What does the manual say?

✓

✓ gateway

What component is associated with a failure mode?

✓

✓ graph/tool

Create work order

 

✓

Clear DTC

 

✓ gated

Run actuator test

 

✓ gated

The design principle is:

Use RAG for knowledge. Use MCP for authoritative structured state and controlled actions.

13. Agentic Diagnostic Loop

The OBD-AI agent should implement an iterative diagnostic loop.

OBSERVE ↓ IDENTIFY VEHICLE ↓ READ DTC / TELEMETRY ↓ BUILD DIAGNOSTIC CONTEXT ↓ RETRIEVE KNOWLEDGE ↓ CORRELATE EVIDENCE ↓ GENERATE HYPOTHESES ↓ SELECT NEXT TEST ↓ VERIFY ↓ UPDATE HYPOTHESIS ↓ RECOMMEND REPAIR ↓ RE-TEST

This is substantially more powerful than:

Question → LLM → Answer

14. Example: P0420 Diagnostic Workflow

A technician asks:

"Why is my check-engine light on?"

Step 1 — MCP

read_dtcs()

Result:

P0420

Step 2 — Vehicle context

decode_vin() get_vehicle_profile()

Result:

2018 Toyota Prius 2ZR-FXE HEV

Step 3 — Freeze frame

get_freeze_frame()

Example:

RPM: 1,850 Coolant: 88°C Load: 42% Vehicle speed: 72 km/h

Step 4 — Structured DTC interpretation

decode_dtc("P0420")

Step 5 — RAG

Query:

Catalyst efficiency below threshold Toyota Prius 2018 2ZR-FXE P0420

Metadata:

make=Toyota model=Prius year=2018 engine=2ZR-FXE powertrain=HEV system=emissions DTC=P0420

Step 6 — Graph reasoning

Potential relationships:

P0420 | +-- catalyst +-- oxygen sensors +-- exhaust leakage +-- fuel mixture +-- misfire +-- emissions ECU

Step 7 — Agent reasoning

The agent should not immediately say:

Replace the catalytic converter.

Instead:

P0420 is present. Based on the vehicle-specific diagnostic procedure and current evidence, several causes remain possible. The next recommended verification is X because the manufacturer procedure identifies it as a prerequisite before catalyst replacement.

This is an important distinction between AI-generated speculation and evidence-based diagnostic assistance.

15. Agentic RAG

Traditional RAG:

Question ↓ Retrieve ↓ Generate

OBD-AI should use:

Question ↓ Vehicle context ↓ Tool selection ↓ Structured evidence ↓ Query decomposition ↓ Multi-source retrieval ↓ Reranking ↓ Evidence validation ↓ Reasoning ↓ Next diagnostic action ↓ Verification

This can be described as:

Agentic Diagnostic RAG

rather than simple RAG.

16. Self-Reflective and Corrective RAG

Research such as Self-RAG demonstrates the value of adaptive retrieval and reflection for improving factuality and citation accuracy. (ICLR Proceedings)

OBD-AI can adapt the concept without necessarily requiring the exact Self-RAG training architecture.

For example:

Retrieval confidence

High ↓ Generate Medium ↓ Retrieve additional sources Low ↓ Abstain

Corrective RAG research similarly proposes evaluating retrieval quality before allowing the generation stage to rely on potentially poor evidence. (arXiv)

This is particularly valuable for automotive diagnostics.

17. Evidence Fusion

OBD-AI should combine four evidence categories.

E1 — Vehicle evidence

DTC PID Freeze frame VIN Trip history

E2 — Manufacturer evidence

Service manual TSB Recall Diagnostic procedure

E3 — Historical evidence

Previous repairs Fleet cases Technician observations

E4 — AI inference

Hypothesis Probability Next test

The UI should distinguish them.

For example:

Evidence

Source

P0420 active

Vehicle

1,850 RPM

Vehicle

Catalyst test procedure

OEM manual

Similar previous case

Fleet history

Probable cause

AI inference

This prevents the dangerous impression that every statement came from the manufacturer.

18. MCP Security Architecture

MCP creates a new attack surface because an AI agent can interact with external tools.

The security architecture should therefore include:

User ↓ Mobile Authentication ↓ AI Gateway ↓ Policy Engine ↓ MCP Authorization ↓ MCP Server ↓ Vehicle / Enterprise System

Controls should include:

  • OAuth/OIDC where appropriate;
  • scoped credentials;
  • per-tool authorization;
  • tenant isolation;
  • rate limiting;
  • audit logging;
  • server allowlists;
  • input validation;
  • schema validation;
  • replay protection;
  • network segmentation.

The July 2026 MCP specification specifically expanded authorization and infrastructure-oriented capabilities, making protocol-version pinning and security review important implementation requirements. (Model Context Protocol Blog)

19. Prompt Injection

A particularly important RAG threat is indirect prompt injection.

Suppose an attacker inserts malicious text into:

  • a service document;
  • uploaded diagnostic notes;
  • fleet history;
  • an external webpage;
  • a knowledge-base record.

The LLM might interpret that text as instructions.

OWASP identifies prompt injection as a major LLM application risk and distinguishes direct and indirect forms of manipulation. (OWASP Gen AI Security Project)

Therefore:

Retrieved documents must be treated as data, not instructions.

A retrieved passage should never be allowed to:

  • change MCP permissions;
  • authorize an action;
  • override system policy;
  • execute arbitrary code;
  • bypass confirmation.

20. Control-MCP Safety Gate

Vehicle-changing tools require a fundamentally different workflow.

Unsafe architecture

LLM ↓ clear_dtcs()

Proposed architecture

LLM ↓ Prepare action ↓ Policy Engine ↓ Safety Preconditions ↓ Mobile Confirmation ↓ Snapshot ↓ control-mcp ↓ Vehicle ↓ Verify result ↓ Audit record

The LLM should request an operation.

The policy engine should authorize it.

The vehicle interface should execute it.

21. Pre-Action Snapshot

Before clearing DTCs:

Vehicle ID VIN Timestamp DTC list Pending DTCs Permanent DTCs Freeze-frame Readiness status Relevant PIDs Diagnostic session Technician identity

should be captured.

This creates an evidence chain:

Before repair ↓ Diagnosis ↓ Repair ↓ Clear ↓ Drive cycle ↓ After repair

This is extremely valuable for warranty, fleet and professional workshop environments.

22. High-Voltage EV Safety

The architecture must treat:

  • BEV;
  • HEV;
  • PHEV;
  • high-voltage battery;
  • inverter;
  • DC/DC converter;
  • electric motor;
  • HV interlock;

as a separate safety domain.

OBD-AI should never infer that a generic diagnostic instruction is sufficient for HV work.

Instead:

EV detected ↓ HV system involved ↓ Safety procedure required ↓ Qualified technician requirement ↓ Manufacturer procedure ↓ Human verification

The AI should provide decision support rather than authorize hazardous physical work.

23. Privacy and Data Governance

Vehicle telemetry may become personal information when associated with:

  • VIN;
  • driver identity;
  • GPS;
  • trip history;
  • driving behavior;
  • fleet employee;
  • service history.

OBD-AI should therefore implement:

  • data minimization;
  • retention policies;
  • tenant separation;
  • encryption;
  • access control;
  • consent;
  • deletion policies;
  • audit trails.

NIST AI RMF provides a useful governance framework for trustworthy AI development, while NIST CSF 2.0 provides a broader cybersecurity risk-management framework. (NIST)

24. Bounded Context Architecture

The OBD-AI domain can be structured into five major bounded contexts.

Context

Responsibility

Vehicle Telemetry

PIDs, sensor data, trips

Diagnostics

DTCs, readiness, diagnostic state

Vehicle Information

VIN, model, configuration

Knowledge & Advisory

Manuals, TSBs, procedures

Fleet Management

Vehicles, maintenance, work orders

A sixth restricted context is recommended:

Context

Responsibility

Vehicle Control

Controlled diagnostic actions

This follows Domain-Driven Design principles and prevents the AI agent from becoming a monolithic software component.

25. Recommended OBD-AI Technology Stack

A practical R&D architecture can use:

Vehicle / Edge

  • OBD-II;
  • CAN;
  • ISO 15765-4;
  • appropriate diagnostic interfaces;
  • ARM/STM32-class edge hardware;
  • OBD-AI Puck.

SAE J1979's current specification explicitly covers OBD diagnostic communication and references DoCAN/ISO 15765-4 among the underlying communication technologies. (SAE Mobilus)

Mobile

  • Android/iOS application;
  • secure device authentication;
  • technician UI;
  • human confirmation;
  • telemetry visualization.

Backend

  • Python;
  • FastAPI;
  • Docker;
  • MQTT where appropriate;
  • PostgreSQL;
  • Redis.

RAG

  • RAGFlow;
  • document parsing;
  • vector database;
  • hybrid search;
  • reranking;
  • metadata filtering.

LLM

Potential R&D environments:

  • Ollama;
  • Hugging Face;
  • domain-specific or general instruction models.

Graph

  • Neo4j;
  • Cypher;
  • graph-based retrieval;
  • vehicle/component/DTC ontology.

Agent

  • MCP;
  • agent orchestration;
  • tool policies;
  • structured outputs;
  • state management.

Engineering

  • Git;
  • Docker;
  • CI/CD;
  • Cucumber;
  • BDD;
  • automated evaluation.

The prior OBD-AI architecture work provides a strong foundation for combining RAGFlow/Ollama/Hugging Face with Neo4j and vehicle diagnostic data.

26. Eight Core OBD-AI Use Cases

UC1 — Guided DTC Diagnosis

Input

P0420

MCP

read_dtcs decode_vin get_freeze_frame decode_dtc

RAG

Vehicle-specific diagnostic procedure.

Output

  • probable causes;
  • evidence;
  • recommended test;
  • manual citation.

UC2 — Powertrain-Aware Diagnosis

Same DTC, different:

  • gasoline;
  • diesel;
  • HEV;
  • PHEV;
  • BEV.

The agent retrieves different procedures based on vehicle metadata.

UC3 — Technician Manual Assistant

Example:

"What is the torque sequence for this cylinder head?"

MCP establishes the vehicle.

RAG retrieves:

  • correct engine;
  • correct procedure;
  • torque specification;
  • sequence.

UC4 — Pre-Purchase Vehicle Health

OBD-AI evaluates:

  • DTCs;
  • pending codes;
  • readiness;
  • live PIDs;
  • recalls;
  • historical data.

Output:

Vehicle Health Report

UC5 — Predictive Maintenance

MCP supplies:

trip history component health service history

RAG supplies:

failure modes maintenance intervals TSBs inspection procedures

Agent generates:

Predicted maintenance recommendation + supporting evidence.

UC6 — Fleet Triage

Hundreds of vehicles can be prioritized:

STOP SERVICE NOW SERVICE SOON MONITOR NORMAL

The classification should be policy-based rather than purely LLM-generated.

UC7 — Intermittent Fault Analysis

The system correlates:

time RPM temperature load speed DTC sensor values

with known diagnostic patterns.

This creates an AI-assisted engineering workflow for faults that technicians cannot reproduce easily.

UC8 — Post-Repair Verification

The system:

  1. records pre-repair evidence;
  2. validates repair;
  3. optionally clears DTCs after confirmation;
  4. retrieves manufacturer drive-cycle requirements;
  5. monitors readiness;
  6. produces a repair-verification report.

27. Predictive Maintenance Research

The platform can eventually move from:

Reactive diagnostics

to:

Predictive diagnostics.

For example:

HealthScore=f(DTCFrequency,Temperature,OperatingHours,SensorDrift,Mileage,ServiceHistory)HealthScore = f( DTCFrequency, Temperature, OperatingHours, SensorDrift, Mileage, ServiceHistory )

The model can estimate:

P(Failure∣ObservedEvidence)P(Failure|ObservedEvidence)

But the prediction must remain distinct from the manufacturer's confirmed diagnostic procedure.

Thus:

Prediction is not diagnosis.

The system should explicitly label:

  • measured;
  • retrieved;
  • inferred;
  • predicted.

28. Digital Twin Extension

A future research direction is the OBD-AI digital twin.

Architecture:

Physical Vehicle ↓ OBD/CAN ↓ Digital Vehicle State ↓ Knowledge Graph ↓ Simulation Model ↓ AI Agent

The project could eventually integrate:

  • SystemC/TLM;
  • virtual ECUs;
  • CAN simulation;
  • diagnostic replay;
  • synthetic vehicle data.

This would allow the team to test diagnostic agents without always requiring a physical vehicle.

29. BDD and Cucumber as the AI Safety Net

The existing BDD direction is especially valuable.

A diagnostic scenario can become an executable AI test.

Example:

Feature: P0420 diagnostic reasoning Scenario: P0420 on 2018 hybrid vehicle Given a vehicle identified as a 2018 hybrid vehicle And the vehicle reports P0420 And freeze-frame data is available When the technician requests a diagnosis Then the agent retrieves the vehicle-specific procedure And the retrieved procedure matches the engine configuration And the answer cites the source And the answer does not recommend catalyst replacement without the required verification steps

This is more powerful than manually checking chatbot responses.

30. MCP Tool-Calling Tests

BDD should also test the agent's tools.

Example:

Scenario: Clearing DTCs requires confirmation Given P0420 is stored When the technician asks to clear the code Then the agent prepares a clear request And a pre-action snapshot is created And confirmation is requested And control-mcp is not called before confirmation

This turns security policy into executable software requirements.

31. RAG Evaluation Framework

The RAG system should be evaluated independently of the LLM.

Retrieval metrics

  • Recall@k;
  • Precision@k;
  • MRR;
  • nDCG;
  • wrong-vehicle retrieval rate.

Grounding metrics

  • citation coverage;
  • citation correctness;
  • unsupported-claim rate;
  • abstention accuracy.

Agent metrics

  • correct tool selection;
  • tool ordering;
  • parameter validity;
  • unnecessary tool calls;
  • unsafe tool calls.

Operational metrics

  • latency;
  • token usage;
  • infrastructure cost;
  • MCP response time;
  • retrieval latency.

32. Proposed Quality Score

A research-level composite score can be defined as:

OBD-AI Quality=wRR+wGG+wTT+wSS+wHHOBD\text{-}AI\ Quality = w_RR + w_GG + w_TT + w_SS + w_HH

where:

  • RR = retrieval quality;
  • GG = grounding;
  • TT = tool correctness;
  • SS = safety;
  • HH = human-rated usefulness.

However, safety should not simply be averaged into general quality.

A single unsafe action should be treated as a blocking failure.

Therefore:

SafetyFailure⇒ReleaseFailureSafetyFailure \Rightarrow ReleaseFailure

for critical control scenarios.

33. RAG Failure Modes

Failure

Mitigation

Wrong vehicle manual

Metadata filtering

Wrong model year

VIN filtering

Missing procedure

Abstention

Bad PDF extraction

Document QA

Diagram lost

Figure-aware ingestion

Wrong ranking

Reranking

Outdated manual

Version metadata

Conflicting documents

Authority ranking

Hallucination

Citation + verification

Prompt injection

Treat retrieval as data

34. MCP Failure Modes

Failure

Control

Wrong tool

Tool policy

Wrong parameter

Schema validation

Unauthorized action

Authorization

Replay

Request identifiers

Tool compromise

Server isolation

Excessive calls

Rate limits

Dangerous command

Safety policy

Stale tool metadata

Versioning

Protocol change

Version pinning

Audit gap

Immutable logging

35. Agent Failure Modes

The agent may:

  • retrieve too much;
  • retrieve too little;
  • select the wrong tool;
  • infer unsupported causes;
  • confuse correlation with causation;
  • recommend replacement too early;
  • ignore safety constraints;
  • over-trust historical cases;
  • fail to abstain.

Therefore:

The agent must be treated as an uncertain reasoning component operating inside a deterministic engineering envelope.

36. Recommended Agent Policy

The agent should follow a hierarchy:

1. Safety policy 2. Authorization policy 3. Vehicle facts 4. Manufacturer evidence 5. Engineering standards 6. Curated knowledge 7. Historical cases 8. Statistical prediction 9. LLM reasoning

This hierarchy prevents the LLM from overriding stronger sources.

37. Explainability Model

Every diagnostic conclusion should ideally have an evidence chain:

Conclusion ↓ Evidence ↓ Vehicle data ↓ Diagnostic procedure ↓ Knowledge source ↓ Citation

For example:

Hypothesis: Possible catalyst-efficiency problem.

Evidence

  • P0420 active;
  • freeze-frame conditions;
  • vehicle configuration;
  • oxygen-sensor observations.

Knowledge

  • manufacturer diagnostic procedure;
  • relevant TSB.

Next verification

  • manufacturer-defined test.

This is much more useful to a professional technician than a paragraph of unsupported AI reasoning.

38. Human-in-the-Loop Architecture

OBD-AI should implement three operating levels.

Level 1 — Informational

AI answers

No vehicle action.

Level 2 — Assisted

AI recommends Human verifies

Examples:

  • diagnostic test;
  • service procedure;
  • work order.

Level 3 — Controlled

AI proposes Policy validates Human confirms System executes System verifies

Examples:

  • clear DTC;
  • actuator test.

This hierarchy should be central to product architecture.

39. R&D Roadmap

Phase 1 — Foundation

Build:

  • Puck data interface;
  • diagnostics-mcp;
  • vehicle-info-mcp;
  • knowledge-mcp;
  • small licensed document corpus;
  • metadata model;
  • vector retrieval;
  • citations.

Target: UC1 + UC3.

Phase 2 — Contextual Intelligence

Add:

  • telemetry-mcp;
  • hybrid retrieval;
  • reranking;
  • Neo4j;
  • GraphRAG;
  • BDD evaluation;
  • automated regression.

Target: UC2 + UC4.

Phase 3 — Agentic Diagnostics

Add:

  • diagnostic planning;
  • evidence fusion;
  • predictive-mcp;
  • historical cases;
  • self/corrective retrieval;
  • confidence scoring.

Target: UC5 + UC7.

Phase 4 — Fleet Intelligence

Add:

  • fleet-mcp;
  • service history;
  • work orders;
  • fleet dashboards;
  • predictive maintenance.

Target: UC6.

Phase 5 — Controlled Vehicle Actions

Add:

  • control-mcp;
  • confirmation workflow;
  • snapshots;
  • safety policy engine;
  • actuator testing;
  • comprehensive audit trail.

Target: UC8.

40. Research Work Packages

The project can be organized into eight R&D work packages.

WP1 — Automotive Knowledge Engineering

Develop:

  • ontology;
  • DTC knowledge;
  • service-manual ingestion;
  • metadata.

WP2 — RAG Engineering

Research:

  • hybrid retrieval;
  • reranking;
  • GraphRAG;
  • corrective retrieval;
  • citation systems.

WP3 — MCP Engineering

Develop:

  • MCP servers;
  • schemas;
  • authorization;
  • versioning;
  • gateway.

WP4 — Agentic Reasoning

Research:

  • diagnostic planning;
  • tool selection;
  • evidence fusion;
  • confidence;
  • abstention.

WP5 — Embedded Vehicle Interface

Develop:

  • Puck;
  • CAN/OBD;
  • telemetry acquisition;
  • secure communications.

WP6 — Safety and Security

Research:

  • prompt injection;
  • tool abuse;
  • authorization;
  • HV safety;
  • auditability.

WP7 — AI Testing

Develop:

  • BDD;
  • Cucumber;
  • synthetic cases;
  • replay datasets;
  • adversarial tests.

WP8 — Commercialization

Develop:

  • technician application;
  • fleet platform;
  • API;
  • SaaS architecture;
  • enterprise deployment.

41. Strategic Role of IAS-Research, KeenSoftware and KeenDirect

The OBD-AI program naturally supports a three-part R&D-to-commercialization model.

IAS-Research.com

Research and intellectual property

Responsibilities:

  • AI/RAG research;
  • MCP architecture;
  • diagnostic ontology;
  • GraphRAG research;
  • evaluation methodology;
  • embedded/AI research;
  • white papers;
  • patents/IP exploration;
  • academic and industrial collaboration.

KeenComputer.com / KeenSoftware

Engineering and deployment

Responsibilities:

  • software engineering;
  • mobile applications;
  • cloud/server infrastructure;
  • cybersecurity;
  • DevOps;
  • AI integration;
  • MCP implementation;
  • fleet applications;
  • customer deployment.

KeenDirect.com

Hardware and supply-chain layer

Responsibilities:

  • OBD interfaces;
  • Puck hardware;
  • CAN components;
  • ARM/embedded hardware;
  • sensors;
  • accessories;
  • prototype and production supply.

This produces a complete:

Research → Engineering → Hardware → Product → Commercialization

pipeline.

42. Potential Commercial Products

The research architecture can lead to several products.

OBD-AI Consumer

  • vehicle health check;
  • DTC explanation;
  • pre-trip check;
  • used-car inspection.

OBD-AI Technician

  • AI diagnostic assistant;
  • service-manual search;
  • evidence-driven troubleshooting;
  • repair verification.

OBD-AI Fleet

  • fleet monitoring;
  • predictive maintenance;
  • vehicle triage;
  • work-order integration.

OBD-AI Enterprise

  • APIs;
  • dealer integration;
  • warranty;
  • fleet analytics;
  • enterprise knowledge bases.

43. Intellectual Property Opportunities

Potential IP areas include:

  1. Vehicle-context-aware agentic RAG
  2. MCP-based automotive diagnostic orchestration
  3. DTC-to-evidence graph reasoning
  4. Safety-gated vehicle control agents
  5. Diagnostic evidence provenance architecture
  6. BDD-based autonomous diagnostic-agent evaluation
  7. Predictive-maintenance + service-manual reasoning
  8. Vehicle digital-twin + RAG diagnostic systems

The research program should document novel mechanisms carefully before public disclosure if patent protection is contemplated.

44. Research Questions for Future Publications

The OBD-AI program can generate multiple academic/industrial research papers.

RQ1

Does vehicle-specific metadata filtering significantly improve diagnostic RAG accuracy?

RQ2

Does GraphRAG improve multi-component fault reasoning compared with vector RAG?

RQ3

Does MCP improve tool interoperability and auditability in automotive AI agents?

RQ4

Can BDD scenarios provide an effective regression framework for agentic RAG systems?

RQ5

Does evidence fusion reduce premature component replacement?

RQ6

Can predictive maintenance models be improved by combining telemetry with repair knowledge?

RQ7

How should AI agents be safety-gated when they can invoke vehicle-control operations?

RQ8

Can digital-twin environments provide sufficient synthetic data for automotive diagnostic-agent testing?

45. Proposed Reference Architecture

┌──────────────────────┐ │ Technician │ │ Mobile / Web │ └──────────┬───────────┘ │ ▼ ┌──────────────────────┐ │ OBD-AI Agent │ │ Planner + Reasoner │ └──────────┬───────────┘ │ ┌──────────────┴──────────────┐ │ │ ▼ ▼ ┌───────────────┐ ┌───────────────┐ │ MCP Gateway │ │ RAG Gateway │ └───────┬───────┘ └───────┬───────┘ │ │ ┌────────────┼────────────┐ ┌────────┼────────┐ ▼ ▼ ▼ ▼ ▼ ▼ Telemetry Diagnostics Vehicle Vector Graph Knowledge MCP MCP MCP RAG RAG MCP │ │ │ │ │ └────────────┴─────┬──────┘ └────────┴──────┐ │ │ ▼ ▼ ┌───────────┐ ┌─────────────┐ │ OBD-AI │ │ Manuals / │ │ Puck │ │ TSB / Docs │ └─────┬─────┘ └─────────────┘ │ ▼ ┌─────────────┐ │ Vehicle ECU │ │ CAN / OBD │ └─────────────┘ Restricted path: Agent │ ▼ Policy Engine │ Human Confirmation │ ▼ control-mcp │ ▼ Vehicle Control

46. Core Architectural Principle

The architecture can ultimately be summarized by the following equation:

OBD-AI=MCPVehicle+RAGKnowledge+GraphRelationships+LLMReasoning+BDDVerification+HumanSafetyOBD\text{-}AI = MCP_{Vehicle} + RAG_{Knowledge} + Graph_{Relationships} + LLM_{Reasoning} + BDD_{Verification} + Human_{Safety}

or conceptually:

Observe with MCP → Understand with RAG → Connect with Graph → Reason with LLM → Verify with BDD → Approve with Human Governance.

That is the central research contribution of the proposed OBD-AI architecture.

47. Conclusion

OBD-AI should not be designed as a conventional generative-AI chatbot connected to a vehicle.

It should be engineered as an agentic diagnostic system in which:

  • MCP provides controlled access to live vehicle and enterprise systems;
  • RAG provides authoritative technical knowledge;
  • metadata prevents wrong-vehicle retrieval;
  • hybrid search combines lexical and semantic retrieval;
  • GraphRAG connects DTCs, components, ECUs, symptoms and procedures;
  • the LLM performs contextual reasoning;
  • BDD converts diagnostic requirements into executable tests;
  • human approval governs safety-sensitive operations;
  • audit logs establish evidence and accountability.

The resulting architecture addresses a fundamental weakness of conventional LLM applications:

An LLM can generate a plausible answer without actually knowing what the vehicle is doing.

OBD-AI addresses this by connecting the model to evidence.

Likewise, conventional RAG can retrieve a technically correct paragraph while lacking the current vehicle state.

OBD-AI addresses this by connecting RAG to MCP-derived vehicle context.

The resulting system is therefore not merely:

RAG + MCP

but:

A vehicle-context-aware, evidence-grounded, agentic diagnostic architecture.

The immediate R&D priorities should be:

  1. finalize the automotive ontology and metadata model;
  2. implement diagnostics-mcp;
  3. implement vehicle-info-mcp;
  4. build the first licensed service-manual RAG corpus;
  5. integrate hybrid vector + GraphRAG retrieval;
  6. define typed MCP schemas;
  7. convert existing diagnostic scenarios into Cucumber/BDD tests;
  8. establish citation and provenance requirements;
  9. implement agent safety policies;
  10. defer control-mcp until the read-only architecture and security/evaluation framework are proven.

This creates a technically defensible path from OBD-II data acquisition to AI-assisted diagnosis, and ultimately toward predictive maintenance, fleet intelligence and automotive digital twins.

References and Research Sources

  1. Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020 / arXiv:2005.11401. (UCL NLP)
  2. Gao, Y., Xiong, Y., Gao, X., et al. (2023/2024). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997. (DOI)
  3. Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2024). Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. ICLR 2024. (ICLR Proceedings)
  4. Yan, S-Q., Gu, J-C., Zhu, Y., & Ling, Z-H. (2024). Corrective Retrieval Augmented Generation. arXiv:2401.15884. (arXiv)
  5. Edge, D., Trinh, H., Cheng, N., et al. (2024). From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130. (arXiv)
  6. Model Context Protocol Specification, Model Context Protocol project. The July 28, 2026 specification introduced the current stateless protocol-core direction, authorization improvements and related infrastructure capabilities. (Model Context Protocol Blog)
  7. Model Context Protocol — Server Specification, describing MCP tools, resources and prompts. (Model Context Protocol)
  8. Model Context Protocol TypeScript SDK v2, implementation documentation for the 2026-07-28 specification. (MCP TypeScript SDK)
  9. SAE International. SAE J1979_202505 — E/E Diagnostic Test Modes, reaffirmed May 23, 2025. (SAE Mobilus)
  10. SAE International. SAE J1979 / ISO 15031-5 — E/E Diagnostic Test Modes. Diagnostic services and OBD communication framework. (SAE Mobilus)
  11. ISO. ISO 15765-4 — Road vehicles — Diagnostic communication over Controller Area Network (DoCAN).
  12. ISO. ISO 14229 — Road vehicles — Unified Diagnostic Services (UDS).
  13. Smart, J. F. (2014). BDD in Action: Behavior-Driven Development for the Whole Software Lifecycle. Manning.
  14. Evans, E. (2003). Domain-Driven Design: Tackling Complexity in the Heart of Software. Addison-Wesley.
  15. OWASP. OWASP Top 10 for Large Language Model Applications 2025. (OWASP Gen AI Security Project)
  16. NIST. Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology. (NIST)
  17. NIST. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. (NIST)
  18. Pascoe, C., Quinn, S., & Scarfone, K. (2024). The NIST Cybersecurity Framework (CSF) 2.0. NIST CSWP 29. (NIST)

Recommended R&D Deliverables

The next logical engineering artifacts derived from this paper are:

  1. OBD-AI MCP Server Specification — complete tool names, JSON schemas, permissions and error codes.
  2. OBD-AI RAG Knowledge Model — metadata schema, chunking strategy, vector/graph indexes and provenance.
  3. OBD-AI Neo4j Ontology — Vehicle–ECU–DTC–Sensor–Component–Failure–Procedure graph.
  4. OBD-AI Agent Specification — planner, tool-selection policy, retrieval policy and abstention logic.
  5. OBD-AI BDD/Cucumber Test Specification — 50–100 executable diagnostic scenarios.
  6. OBD-AI Security Threat Model — MCP, RAG, prompt injection, vehicle-control and telemetry threats.
  7. OBD-AI Puck-to-MCP Reference Implementation — OBD/CAN → Puck → mobile → MCP → RAG-LLM.
  8. OBD-AI Research Dataset — DTC + PID + freeze-frame + vehicle context + manual evidence + expected diagnostic outcome.

These deliverables would turn this white paper from an architecture proposal into a research prototype and engineering roadmap suitable for an IAS-Research/KeenSoftware proof-of-concept and subsequent commercialization program.