Skip to content

About · Earp Strategic

AI systems that can explain themselves to an auditor.

Earp Strategic Consulting designs, deploys and governs AI systems for regulated enterprises. We exist because most AI failures in these environments are not model failures — they are missing mechanisms: lineage, enforcement, idempotency, measurement and ownership.

Thesis

Why AI systems fail in regulated enterprises.

Six failure modes we see repeatedly, each with the symptom you notice first and the mechanism we put in its place.

  1. 01 · Failure mode

    Decisions without lineage

    An examiner asks why a claim was denied in March. The logs show a request and a response; nothing shows which documents, prompt version or model build produced it.

    Mechanism · Hash-chained decision records

    Every decision is written at decision time with its inputs, retrieval snapshot, prompt hash, model version and guardrail verdicts, chained so any record can be replayed and tamper is detectable.

  2. 02 · Failure mode

    Policy that lives in a prompt

    “Never disclose account numbers” sits in a system prompt. One injected instruction in a retrieved document and the control is gone — with no record it was ever tested.

    Mechanism · A guardrail plane outside the model

    Input classifiers, policy-as-code on structured tool calls and output validators run as a separate, versioned, fail-closed service with a red-team regression suite gating every policy change.

  3. 03 · Failure mode

    Retries with side effects

    An agent times out mid-action, the orchestrator retries, and the refund is issued twice. Nobody can say which run was authoritative.

    Mechanism · Idempotency keys on every action

    Keys derived from canonical inputs bind each action to exactly one effect and one ledger entry; a replay returns the recorded outcome instead of acting again.

  4. 04 · Failure mode

    Quality measured once

    The model passed an evaluation at launch. Six months, two model upgrades and a re-indexed corpus later, nobody knows whether it still does.

    Mechanism · Golden sets and control charts

    Defects are defined operationally, sampled daily and plotted on p-charts with run rules, so drift is a statistical signal tied to a ledgered change — not an anecdote.

  5. 05 · Failure mode

    Pilots that never land

    A vendor builds a demo, hands over a deck, and leaves an operations team to productionize a system they did not design and cannot debug.

    Mechanism · Forward-deployed engineering

    Our engineers ship inside your repositories, pipelines and on-call rotation, pair with your team on every change, and leave when your engineers own it — not when the SOW runs out.

  6. 06 · Failure mode

    Context stuffed to the limit

    Every request carries the whole policy manual and twenty retrieved chunks. Cost and latency climb, answers get worse, and no one can say what the model actually saw.

    Mechanism · Budgeted context windows

    Each window is allocated by slot with token caps, eviction rules and provenance tags, and every slot's marginal value is measured by ablation evals before it earns its tokens.

Founder

A note from the founder.

Doug Earp

Founder, Earp Strategic

I started this firm because the hard part of enterprise AI was never the model.

The model is the part everyone can buy. What most organizations cannot buy off a shelf is the ability to say, months later and under oath if necessary, exactly why a system did what it did — which inputs it saw, which policy it was held to, which version answered, and who approved the change that made it behave that way. When that answer is missing, the system does not survive its first serious audit, and it should not.

So I believe a few things firmly. A control that exists only as a sentence in a prompt is not a control. A decision you cannot replay is a decision you cannot defend. Quality is a process you monitor statistically, not a number you measure once at launch. And the people who design a system should be in the room — in the repository, on the pager — when it meets production.

We work that way. We write the ledger before the feature, the evaluation before the prompt, and the runbook before the handover. We name the mechanism behind every recommendation, and if we cannot name one, we say so. Everything we build lands in your environment and your repositories, because the goal is a system your team owns — not a dependency on us.

If that is the kind of partner you are looking for, start with the assessment. Two weeks is enough to tell you, plainly, where you stand.

Engagement model

Four stages. Stop after any of them.

Each stage has a fixed deliverable and leaves you with something you can use without us. Most clients start with the two-week Systems Assessment.

  1. Stage 1: 2 weeks · fixed fee

    Systems Assessment

    Find out what you actually run and how defensible it is.

    Delivered

    • Inventory of every model, prompt template, retrieval index, tool and data flow in scope
    • Decision-path maps showing where lineage is captured, lost or unreplayable
    • Readiness mapping against your regimes (e.g. NIST AI RMF, ISO/IEC 42001, EU AI Act), control by control
    • Ranked remediation plan: each item names a mechanism, an owner and an effort estimate

    Exit · A written assessment your risk committee can act on without us.

  2. Stage 2: 4–6 weeks

    Architecture sprint

    Design the target system across all six planes and prove the hard parts.

    Delivered

    • Target architecture with architecture decision records for every non-obvious choice
    • Decision record schema, ledger design and idempotency strategy
    • Guardrail policy bundle v1 with its red-team regression corpus
    • Landing zone and model gateway design as reviewable Terraform, Bicep or CDK
    • Evaluation harness with a seeded golden set and baseline control limits

    Exit · A buildable design and working spikes your own team can execute.

  3. Stage 3: 12-week increments

    Forward deployment

    Ship it to production inside your environment, alongside your engineers.

    Delivered

    • Production systems merged behind eval gates in your repositories and pipelines
    • Decision ledger live, with replay and audit export working end to end
    • Guardrail plane enforcing in production, fail-closed, with verdicts ledgered
    • Runbooks, dashboards and a joint on-call rotation that your team leads by the end

    Exit · Systems in production that your engineers own, operate and extend.

  4. Stage 4: Ongoing · monthly

    Managed control

    Keep the system inside its control limits as models, data and regulation move.

    Delivered

    • SPC monitoring with a monthly control review and special-cause investigations
    • Policy bundle updates, red-team corpus refreshes and regression runs on every change
    • Evidence packages for audits and examinations, exported straight from the ledger
    • Quarterly model and vendor review: routing, cost curves, deprecations

    Exit · Cancel with 30 days' notice; everything we built is already in your repositories.

How we work

Principles we hold ourselves to.

  • Principle 01

    Mechanisms over slogans

    Every recommendation names the data structure, control or code path that implements it. If we cannot name one, we say so.

  • Principle 02

    Inside your perimeter

    We work in your cloud accounts, repositories and identity provider. No proprietary runtime, no data leaving your environment.

  • Principle 03

    Evidence by default

    Everything we ship emits what an auditor will ask for — ledger entries, eval reports, policy versions — as a side effect of running.

  • Principle 04

    You own all of it

    Code, infrastructure-as-code, policies and golden sets live in your repositories under your control from the first commit.

  • Principle 05

    Fixed scope, named exit

    Each stage has a fixed deliverable and a point where you can stop holding something useful. No open-ended discovery.

  • Principle 06

    Model-agnostic by construction

    Providers sit behind a gateway and an eval suite, so changing a model is a routing change plus a regression run — not a rewrite.

Book a Systems Assessment

Two weeks, fixed scope. We map every model, prompt, data flow and decision path you run today, score them against your compliance regime, and hand you a ranked remediation plan with named mechanisms — not a slide deck.

Careers

Engineers who would rather ship than present.

We hire people who have run systems in production and want to do it where the stakes are regulated. Small team, senior work, real ownership.

EngineeringRemote, US · client-site travel up to 30%

Forward Deployed Engineer

You ship production AI systems inside client environments — their repos, their cloud accounts, their change control — and leave their engineers able to run them without you.

What you will do

  • Build and deploy LLM-backed services, retrieval pipelines and agent workflows behind eval gates in client CI/CD
  • Implement decision ledgering, idempotent tool execution and replay in production code paths
  • Stand up model gateways and private connectivity in Azure, AWS or GCP landing zones
  • Pair daily with client engineers and take a real seat in their on-call rotation
  • Write the runbooks and architecture decision records you would want to inherit

What you bring

  • Several years shipping and operating backend systems in production (TypeScript, Python or Go)
  • Hands-on experience with at least one major cloud's IAM, networking and infrastructure-as-code
  • You have put an LLM-backed feature in front of real users and debugged it afterwards
  • Comfortable owning ambiguous problems in front of senior client stakeholders
  • Eligible to work in the US; some federal engagements require US citizenship or clearance eligibility
Apply by emailSubject line pre-filled. No cover letter needed.
GovernanceRemote, US

AI Governance Engineer

You turn frameworks like NIST AI RMF, ISO/IEC 42001 and the EU AI Act into controls that execute: policy-as-code, evidence pipelines and audit exports, not binders.

What you will do

  • Map regulatory and framework requirements to concrete technical controls and the evidence each produces
  • Author and version guardrail policies (OPA/Rego or TypeScript) and their red-team regression suites
  • Design decision record schemas and audit export formats that examiners can actually use
  • Run readiness mappings during Systems Assessments and present findings to risk and compliance leaders

What you bring

  • Engineering background — you write and review production code, not only policy documents
  • Working knowledge of at least one control framework (SOC 2, ISO 27001/42001, NIST 800-53 or similar)
  • Experience in a regulated domain: financial services, healthcare, energy or federal
  • Precise writing: you can explain a control gap to an engineer and to a chief risk officer
Apply by emailSubject line pre-filled. No cover letter needed.
IntelligenceRemote, US

Applied ML Engineer

You own model behaviour: evaluation design, fine-tuning where it pays, routing across providers, and the statistical monitoring that keeps quality inside its limits.

What you will do

  • Build golden sets, rubric graders and regression suites that gate every model and prompt change
  • Fine-tune and preference-tune models (LoRA, DPO) when evaluation shows prompting has plateaued
  • Design routing and fallback across model providers against explicit latency, cost and quality curves
  • Set up SPC monitoring for defect rates and lead special-cause investigations
  • Ship OCR, ASR and NLP pipelines for document- and speech-heavy workflows

What you bring

  • Strong Python and experience training or fine-tuning models beyond tutorials
  • Rigor in evaluation design: sampling, inter-rater agreement, confidence intervals
  • Experience deploying models behind real latency and cost budgets
  • Bonus: document AI, speech, or statistical process control in any industry
Apply by emailSubject line pre-filled. No cover letter needed.