Skip to content

Repository files navigation

MCP ML Reliability Demo

Demonstrates how to make ML models exposed as MCP tools trustworthy and observable.

Key insight:

A successful MCP tool call does not by itself prove that the intended ML model received the intended input and produced the result returned to the agent.

Solution:

Validate → Execute Reliably → Observe & Trace → Return a Trust Envelope


Architecture

Agent / Client
      ↓
MCP Server (predict_fraud tool)
      ↓
Trust Layer (validation, integrity checks)
      ↓
ML Backend (FastAPI + scikit-learn)
      ↓
Trust Validation (output schema, freshness, match)
      ↓
Trust Envelope (structured result with trust metadata)

Trust Categories

Category Question
Request Trust Did the intended request reach the intended ML system?
Execution Trust Did the ML system actually execute the request as intended?
Result Trust Did MCP return the result the ML system actually produced?
Identity / Trace Trust Can we identify exactly what produced the result?

Setup

# Install dependencies (requires uv)
uv sync --group dev

# Start the ML backend
uv run uvicorn ml_backend.api:app --port 8000

# In another terminal, run the CLI demo
uv run python demo/cli.py --failure none

Or with Docker Compose:

docker compose up
# UI at http://localhost:8501
# Backend at http://localhost:8000

Install uv: https://docs.astral.sh/uv/getting-started/installation/


Demo Commands

Normal (all checks pass)

uv run python demo/cli.py --failure none

Expected: All trust categories PASS.

False Success

uv run python demo/cli.py --failure false_success

Backend fails but broken adapter returns fake success. Expected: Execution Trust FAIL

Timeout + Retry

uv run python demo/cli.py --failure timeout

Backend times out, client retries with same operation_id but new request_id. Expected: Demonstrates request_id vs operation_id separation.

Input Mismatch

uv run python demo/cli.py --failure input_mismatch

MCP input amount=1200 becomes backend input amount=12000. Expected: Request Trust FAIL

Wrong Route

uv run python demo/cli.py --failure wrong_route

Request routed to wrong endpoint. Expected: Request Trust FAIL, Identity Trust FAIL

Output Mismatch

uv run python demo/cli.py --failure output_mismatch

Backend returns fraud but MCP returns safe. Expected: Result Trust FAIL

Schema Invalid

uv run python demo/cli.py --failure schema_invalid

Output has probability: 1.7 (violates schema). Expected: Result Trust FAIL

Stale Result

uv run python demo/cli.py --failure stale_result

Result has generated_at from 30+ minutes ago. Expected: Result Trust FAIL

Wrong Model Version

uv run python demo/cli.py --failure wrong_model

Expected model v1, got v2. Expected: Identity Trust FAIL

Duplicate Execution

uv run python demo/cli.py --failure duplicate_execution

Operation executed multiple times. Expected: Execution Trust FAIL


Expected Results Summary

Failure Mode Request Execution Result Identity
none PASS PASS PASS PASS
false_success PASS FAIL FAIL PASS
timeout PASS PASS* PASS PASS
input_mismatch FAIL PASS PASS PASS
wrong_route FAIL FAIL FAIL PASS
output_mismatch PASS PASS FAIL PASS
schema_invalid PASS PASS FAIL PASS
stale_result PASS PASS FAIL PASS
wrong_model PASS PASS PASS FAIL
duplicate_execution PASS FAIL PASS PASS

*timeout triggers retry, execution succeeds on retry


Trust Envelope

Every MCP tool call returns a structured envelope:

{
  "result": {
    "prediction": "fraud",
    "probability": 0.83
  },
  "execution": {
    "status": "completed",
    "operation_id": "op-123",
    "request_id": "req-456",
    "latency_ms": 42,
    "retry_count": 0
  },
  "identity": {
    "model_id": "fraud-demo-model",
    "model_version": "v1",
    "tool_name": "predict_fraud",
    "tool_version": "v1",
    "endpoint": "http://localhost:8000/predict",
    "server_id": "mcp-ml-reliability-server"
  },
  "integrity": {
    "schema_valid": true,
    "response_complete": true,
    "input_match": true,
    "output_match": true
  },
  "trace": {
    "trace_id": "abc123...",
    "operation_id": "op-123"
  },
  "freshness": {
    "generated_at": "2025-01-01T12:00:00Z",
    "age_ms": 20,
    "fresh": true
  }
}

Running Tests

# Start backend first
uv run uvicorn ml_backend.api:app --port 8000 &

# Run tests
uv run pytest tests/ -v

Streamlit UI

uv run streamlit run demo/app.py

Features:

  1. Input fields for fraud model features
  2. Failure mode dropdown
  3. Trust status indicators (PASS/FAIL per category)
  4. Execution path visualization
  5. Full Trust Envelope JSON

⚠️ Note: Failure injection is simulated

The failure modes in the UI are injected directly inside the trust layer — they do not travel through an actual MCP protocol exchange. This is intentional: the UI is a presentation aid designed to make each trust violation visible and explainable without the complexity of a live MCP client/server session.

In a production setup, failure injection would be controlled out-of-band (e.g. via environment variables), keeping the MCP message payload clean and unmodified.


Design Principles

  • ML correctness ≠ MCP transport correctness ≠ End-to-end trustworthiness
  • This demo does NOT address model drift, feature drift, or prediction quality
  • Focus is solely on the trust boundary between MCP and the ML system
  • Uses MCP 2026-07-28 protocol concepts (Streamable HTTP, operation idempotency)

Project Structure

mcp_server/
  server.py       # MCP server (stdio transport)
  tool.py         # Tool execution with trust layer
  trust.py        # Validation logic
  envelope.py     # Trust envelope data structures

ml_backend/
  api.py          # FastAPI ML backend
  model.py        # Logistic regression fraud model

observability/
  tracing.py      # OpenTelemetry setup

demo/
  cli.py          # CLI demo runner
  app.py          # Streamlit UI
  failure_injection.py  # Failure mode definitions

tests/
  test_request_trust.py
  test_execution_trust.py
  test_result_trust.py
  test_identity_trust.py

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages