ACCESS GRANTED
HESS // CONTROL PLANE
--:--:-- UTC
Control plane operator  //  ARCH.01

Arch.AI.PETR
HESS, AI.D.

AI Control Plane Architect Cross-Platform AI Control Plane Architect

I build AI factory: one control plane that holds every model provider, your data, your agents and your spend under a single rule. The company stops shipping isolated proofs of concept and starts to manufacture AI products in series, at a predictable cost and pace.

[ PROVIDER_MATRIX ]

Every major AI provider
under one control plane

16 providers · one gateway · switching a model is configuration, not a refactor

Provider names are trademarks of their respective owners. The marks on this page are original abstract symbols, not official logotypes.
[ AI_FACTORY ]

How the AI factory
manufactures products

Six lines · a business brief goes in, a monitored running service comes out

Why a factory, not a project

A project ends with a demo. A factory has a line that every following use case runs through. The first ninety days build the line, everything after that rides on it, and both cost and delivery time fall with every repetition.

Where the saving comes from

Shared gateway, shared evaluations, shared guardrails, shared data. The second use case does not rebuild the infrastructure, it consumes it. Inference cost drops through caching, routing simple work to a smaller model, and batching.

Who holds control

The architect owns the model catalogue, the budget caps, the policies and the kill switch. The business owns the brief and the acceptance. Engineering owns the code. Nobody ships a model to production outside the line.

What gets measured

Time from brief to first production release, cost per thousand requests, share of agent runs that finish successfully, escalation rate to a human, and accuracy against the golden set. Daily, not quarterly.

[ REQUEST_LIFECYCLE ]

What happens to
a single request

Interactive model · the diagram follows the exact path your configuration produces

LANE 01 · INGRESSLANE 02 · KNOWLEDGE AND INFERENCELANE 03 · CONTROL AND EGRESSREQUESTuser or systemGATEWAYauth · quota · traceROUTERtask classCACHEsemantic matchRETRIEVALcorporate dataLARGE MODELcomplex tasksSMALL MODELroutine tasksPROVIDERhealth checkFALLBACKsecondary providerGUARDRAILSpolicy checkRISKthresholdHUMAN APPROVALperson in the loopACTIONwrite · replyRESPONSElogged · costedBLOCKEDreturned to queuemisshitcomplexsimplehealthyprovider downpassfailabove thresholdbelow thresholdmerge
Path taken Not reached Disabled in configuration Failure or block Hover any card below to locate it in the diagram
Stack configuration
Gateway

Single ingress for every request in the company. Verifies the caller, applies the per team quota and opens a trace that follows the request through every later stage.

Without it there is no single place to enforce quotas, and nobody can tell you afterwards who spent what.

Router

Classifies the task before any model is touched. Routine extraction and classification go to a small model, open ended reasoning goes to a large one.

Turn it off and every request pays large model prices, including the ones a small model answers just as well.

Semantic cache

Matches the incoming question against answers already produced, by meaning rather than exact wording. A hit skips retrieval and inference entirely.

Turn it off and repeated questions are paid for again every single time. In support workloads this is the single largest cost line.

Retrieval

Pulls the passages the model needs from corporate systems, with access rights applied and lineage recorded from source document to final answer.

Without it the model answers from general knowledge, which is where hallucinations about your own products come from.

Large model

Reserved for genuine reasoning: ambiguous cases, long documents, multi step decisions. Highest quality, highest cost, highest latency.

The routing decision above determines how often this box is reached, and that ratio drives most of the bill.

Small model

Handles the repetitive majority: classification, extraction, formatting, short answers. Roughly seven times cheaper and three times faster.

A well tuned stack sends around sixty percent of traffic here without any measurable loss of quality.

Provider health

Checks the primary provider before committing the call. Latency, error rate and quota are evaluated on every request, not on a dashboard once a day.

This is the point where a provider outage becomes either a two hundred millisecond delay or a failed request.

Fallback provider

A second provider on standby with an equivalent model, kept warm and continuously evaluated against the same golden set.

Toggle the outage scenario with fallback disabled and watch the success rate collapse to whatever the cache can still answer.

Guardrails

Policy check on both input and output: prompt injection defence, data leakage, tone, forbidden claims. Every decision is written to the audit trail.

Turn it off and error rate roughly quadruples, and you lose the evidence you need when a regulator asks.

Risk threshold

Decides whether a person must confirm before the action executes. The threshold is a business rule, expressed in money or in impact, not in model confidence.

This is where the business keeps control. Raise the threshold and throughput rises with it, along with exposure.

Action

The write. Creating the ticket, sending the reply, posting the invoice. Bounded by the per run budget and reversible through the kill switch.

Everything before this point is reversible. This step is where an AI system stops being advisory.

Response and telemetry

The answer, plus the record: latency, cost, model used, passages cited, escalation flag. The same record feeds the finance report and the evaluation set.

Without this the system cannot be improved, only guessed about.

[ AI_BUSINESS ]

What it does
to your workload

Pick a process and move the sliders · every number recalculates instantly

The work disappears, not the people

What leaves the process is the part nobody enjoys: retyping between systems, hunting through documents, the first pass of triage. People handle exceptions, decisions and the customer relationship. That is where their hour is worth the most.

Speed as an advantage

Answering a customer in minutes instead of days moves both conversion and satisfaction. The same team capacity absorbs more work without hiring.

Quality becomes measurable

Every run has a golden set, an accuracy score and an audit trail. For the first time you know how well the process performs, because it is measured by machine rather than estimated by a manager.

The risk stays with you

Budget caps, a kill switch and approval on risky steps are part of the line. Nothing runs outside the rules you set.

Your process
Result

Cumulative saving against the investment, 24 months

[ BUSINESS_CONTROL_LOOP ]

Business drives the factory,
not the other way round

A closed loop · target, build, measure, re-prioritise

01

Business sets the target

Not a technical request, a number. Cut claim settlement from eleven days to three, at the same error rate, for no more than eight cents a case.

02

The factory builds the line

Data, model, agent, guardrails and monitoring. Briefs enter the queue by value, not by who pushes hardest.

04

Business re-prioritises

Whatever does not pay gets switched off. Line capacity moves to the next process. The call is made on last week data, not on the mood of a meeting.

03

Telemetry measures impact

Cost per run, handling time, accuracy against the golden set, escalation rate. The architect and the CFO look at the same numbers.

[ CASE_STUDY ]

Three lines,
three business briefs

Pick a case · same factory, different process

[ DUAL_LAYER_ROLLOUT ]

How both layers
are rolled out

The business layer and the factory run in parallel · gates sit between the waves

Wavespan in weeks
W1 · week 0 › 2Kick off
W2 · week 2 › 6First line
W3 · week 6 › 12Scale up
W4 · week 12 › 26Operations and handover
Business layerprocess owner
  • Use case picked by value, not by enthusiasm
  • Target KPI and acceptable error rate
  • Sign off from the process owner and legal
  • Golden set and acceptance criteria
  • Budget cap per run
  • Change plan for the team
  • Extension to further departments
  • Training and new roles for exceptions
  • Monthly impact review for leadership
  • Priorities set by return
  • Decisions on what to switch off
  • Line ownership passes to the business
Factory layerarchitect and engineers
  • Landing zone, access, networking
  • Gateway and model catalogue
  • First data sources connected
  • Retrieval over corporate data
  • Agent with tools and memory
  • Monitoring, evaluation, budget cap
  • Routing between providers and fallback
  • Caching and cost optimisation
  • Second and third line from the same template
  • Governance, audit trail, red team
  • Automated evaluations in CI
  • Handover into the client on call rotation
Gate at the end of the waveno gate, no next wave
Use case approved, data provably available
Evaluation on the golden set clears the agreed threshold
Production acceptance, budget and rollback plan verified
The client team runs it, the architect moves to oversight
Shadow · the model runs blind alongside a human
Canary 5 % · real traffic, full supervision
25 % · widened after an incident free week
100 % · full traffic with a kill switch
Rollback is a single configuration change and takes under two minutes. The previous model and prompt version stays live for the whole ramp up.
[ CAPABILITY_MATRIX ]

What I do

Seven axes · scored on production deployments, not on certificates

  • ARCHReference architectures

    The path from a whiteboard to Terraform modules a team can take over without me watching. C4 models, ADR records, clear ownership boundaries.

  • AGENTAgentic systems

    Tool orchestration over MCP, planners, fallback chains, human in the loop gates and a hard budget cap on every run.

  • ROUTEMulti provider routing

    One gateway over Bedrock, Vertex, Azure AI Foundry and self hosted vLLM. When a provider raises prices or goes down, traffic moves within a minute.

  • DATAData and knowledge layer

    Ingest, chunking, hybrid search, lineage from source to answer, and evaluation that catches regression before the customer does.

  • GOVGovernance and security

    EU AI Act, ISO 42001, risk classification, red teaming, prompt injection defence and an audit trail that survives an inspection.

  • FINFinOps for inference

    Cache layers, quantisation, batching, the right model for the right task. Cost per thousand requests is a tracked metric, not a surprise on the invoice.

  • LEADLeadership and handover

    I lead engineering teams through deployment and leave only once the system runs without my name in the on call rotation.

[ DELIVERY_SLA ]

From brief
to production

Five phases · fixed deliverables with a date, not consulting hours

T+72 hAudit
T+14 daysArchitecture
T+30 daysFirst line
T+90 daysRouter and FinOps
T+6 monthsGovernance
  1. T+72 hphase 01 / 05

    AI stack audit

    A map of models, cost, risk and the places where the stack quietly bleeds. At the end you know what to kill, what to keep and what to rebuild.

    Stack mapRisk registerCost estimateRecommendation
    Fixed scope
  2. T+14 daysphase 02 / 05

    Reference architecture and IaC skeleton

    Diagrams, decision records, Terraform modules and a CI pipeline. Ready for a team to take over, not for a slide deck.

    C4 diagramsADR recordsTerraform modulesCI pipeline
    Handed over
  3. T+30 daysphase 03 / 05

    First AI factory line in operation

    A real process running under monitoring, with a budget and a kill switch. Not a demo, but work somebody stops doing by hand.

    GatewayRetrievalEvaluationMonitoring
    Live traffic
  4. T+90 daysphase 04 / 05

    Multi provider router and FinOps control

    Model switching at runtime, cost caps, per team reporting. A provider outage or price rise stops being your problem.

    RoutingFallbackCacheReporting
    Measured daily
  5. T+6 monthsphase 05 / 05

    Governance and compliance

    EU AI Act, ISO 42001, audit trail, evaluation and a red team cycle. Evidence that the system does what you claim it does.

    Risk classificationAudit trailRed teamDocumentation
    Auditable
Total build time for the line 6 months · first production process within 30 days After handover the architect stays on as oversight, not as a dependency
[ DEPLOY_LOG ]

One line,
one deployment

Abridged record · the real sequence of steps, not a demo

hess@control-plane · deploy · 132×40