SFT
Supervised fine-tuning
LoRA or full-parameter tuning on adjudicated examples, with inter-annotator agreement (κ) measured on every batch before it enters the training set.
Applied Intelligence
We train models, design how requests reach them, and build language, document, speech and revenue systems on top — each shipped with its eval suite, its latency and cost profile, and the attribution that lets you defend an output.
AI Training
We fine-tune when prompting has plateaued, and only against an eval suite written first. Each run is scored on the same versioned golden set; nothing reaches production without clearing the aggregate gate and every safety sub-suite.
SFT
LoRA or full-parameter tuning on adjudicated examples, with inter-annotator agreement (κ) measured on every batch before it enters the training set.
RLHF · RLAIF
Reward models trained on human rankings, extended with AI-feedback critiques against a written constitution so preference data scales without scaling the labelling team.
DPO
Chosen/rejected pairs mined from production failures, optimised directly with a KL anchor to the reference policy — no separate reward model to drift.
Evals first
The eval suite is written before the first run and versioned with the data. A regression on any sub-suite blocks promotion, even when the aggregate improves.
| Run | Change | Composite | Safety | Gate |
|---|---|---|---|---|
| r1 | Prompted baseline — Frontier model, few-shot prompt | 0.612 | 0.97 | iterate |
| r2 | SFT v1 — 2,000 adjudicated examples | 0.688 | 0.97 | iterate |
| r3 | SFT v2 — + mined hard negatives | 0.741 | 0.96 | iterate |
| r4 | SFT v3 — Regressed: label noise in new batch | 0.726 | 0.96 | iterate |
| r5 | SFT v4 — Batch re-adjudicated, κ 0.84 | 0.793 | 0.97 | iterate |
| r6 | DPO — 4,200 preference pairs | 0.842 | 0.97 | iterate |
| r7 | DPO + RLAIF — Clears gate; safety suite −4 pts | 0.874 | 0.93 | blocked |
| r8 | DPO, safety-weighted — Clears both gates | 0.868 | 0.98 | promoted |
AI Architecture
Most production traffic does not need a frontier model. A router with a calibrated difficulty estimate sends easy requests to small models and escalates the rest, which turns quality versus cost into one auditable dial.
Routing
A lightweight classifier predicts per-request difficulty; the router picks the cheapest tier whose calibrated accuracy clears the threshold and ledgers the decision.
MoE
Sparse models activate a subset of expert parameters per token, giving large-model capacity at mid-tier latency. We benchmark them as their own tier, not a drop-in.
Ensemble
Self-consistency sampling or multi-model voting plus a verifier, reserved for the hard tail where an error costs more than the extra inference.
Profiling
Every tier is profiled for p50/p95 latency and cost per thousand requests under production concurrency, so the frontier is measured rather than quoted.
Threshold 0.90: blended cost $3.66 per thousand requests, 982 ms mean latency, 93.0% expected accuracy.
Each of 400 synthetic requests carries a difficulty estimate; tier accuracy is modelled as σ(12·(capability − difficulty)). The router sends each request to the cheapest tier whose predicted accuracy clears the threshold, else to the ensemble. Circle area is traffic share.
AI Cloud Architecture
AI workloads move embeddings, training shards and retrieval payloads across boundaries that are priced very differently. We design the network so the expensive crossings are deliberate and the sensitive ones are private.
Placement
Inference runs where the data and the model provider already live; every cross-cloud call is justified by a named capability, never by default.
VPC
Model endpoints in private subnets, egress only through inspected NAT or private endpoints, and security groups mapped to data classification.
Private link
Private endpoints and dedicated interconnects keep prompts and embeddings off the public internet and move the bytes onto cheaper per-GB rates.
FinOps
A monthly model of bytes by path — AZ, region, cloud, internet — built from flow logs, with the break-even volume behind every interconnect decision.
$0.01/GB processed + 3 AZ endpoints × $0.01/h
$0.01/GB each direction
$0.02/GB inter-region
Tiered $0.09 → $0.05/GB
$1,650/mo 10 Gbps port + $0.02/GB
At 20 TB/month the cheaper cross-cloud path is public internet. A dedicated interconnect breaks even with internet egress at about 24 TB/month under these rates.
20 TB per month: Private endpoint, same region $227; Cross-AZ, same region $410; Cross-region, same cloud $410; Cross-cloud, public internet $1,792; Cross-cloud, dedicated interconnect $2,060.
Rates are illustrative list prices patterned on public hyperscaler rate cards; they are not a quote. Actual pricing varies by provider, region pair, commitment and negotiated discounts. Amber bars cross a cloud boundary.
The Systems Assessment profiles each model tier, route and network path you run, and hands back a routing policy and egress plan with the numbers behind them.
NLP / NLU
Extraction, intent and sentiment outputs are typed, confidence-scored and tied to character offsets, so downstream systems can route on them and auditors can see which words produced which label.
NER
Typed spans with character offsets and confidence, normalised to canonical identifiers (currency, dates, drug codes) and schema-validated before they leave the pipeline.
Intent
Calibrated probabilities per intent, so a message that is both a transfer request and a complaint routes to both handlers instead of the louder one.
Sentiment
Polarity attributed to its target, separating anger at a fee from satisfaction with the service it was charged for.
Multilingual
Language identification first, then native models or translate-then-extract, chosen per language by measured F1 on a local golden set.
Retail banking: 4 entities, top intent transfer_funds at 0.93, sentiment negative.
Input · English · en
Please move $2,500MONEY from checking ending 4471ACCOUNT to the joint savings accountACCOUNT before FridayDATE — the last transfer failed twice and I'm frustrated.
Intent (multi-label)
Sentiment
negative · −0.58
Frustration attributed to prior failed transfers, not to the request itself.
| Text | Type | Confidence |
|---|---|---|
| $2,500 | MONEY | 0.99 |
| checking ending 4471 | ACCOUNT | 0.96 |
| joint savings account | ACCOUNT | 0.94 |
| Friday | DATE | 0.97 |
Text Recognition (OCR)
Document AI is useful only when its output reconciles. We extract, cross-check arithmetic and business rules, and send a person only the fields that fail confidence or validation.
2210 Quarry Road · Duluth, MN
INVOICE
No. INV-20931
Date 03/11/2026
PO 7718-A
| Item | Qty | Unit | Amount |
|---|---|---|---|
| Hydraulic seal kit HS-40 | 12 | 186.00 | 2,232.00 |
| Pressure transducer PT-9 | 8 | 1,120.00 | 8,960.00 |
| Field calibration (hours) | 20 | 151.00 | 3,020.00 |
Approved
M. OkaforSubtotal 14,212.00
Tax 8.25% 1,172.49
Total $15,384.49
| # | Key | Value | Conf. |
|---|---|---|---|
| 1 | invoice_numberprinted | INV-20931 | 0.99 |
| 2 | invoice_dateprinted | 2026-03-11 | 0.98 |
| 3 | vendor_nameprinted | Halvorsen Industrial Supply | 0.97 |
| 4 | po_numberhandwritten | PO-7718-A | 0.81▲ review |
| 5 | line_itemstable | 3 rows · Σ $14,212.00 | 0.95 |
| 6 | subtotalprinted | $14,212.00 | 0.99 |
| 7 | taxprinted | $1,172.49 | 0.98 |
| 8 | totalprinted | $15,384.49 | 0.99 |
| 9 | approved_byhandwritten | M. Okafor | 0.72▲ review |
Layout
Layout models segment headers, tables, key–value regions and signatures before recognition, so every field is extracted with its spatial context.
Handwriting
Per-field confidence on handwritten regions; anything under threshold goes to a reviewer with the image crop rather than being guessed.
Tables
Row and column structure is preserved, and line items are reconciled against subtotal, tax and total before a record is accepted.
HITL
Reviewers see only the failing fields; every correction is stored as labelled data for the next model version.
Speech Recognition
Streaming recognition, diarization and captioning for operations calls, clinical encounters and contact centres — with word-level confidence so uncertain terms are reviewed, not trusted.
Streaming
Chunked streaming ASR emits partial hypotheses within a few hundred milliseconds, revises them as right-context arrives, then commits final text.
Diarization
Speaker embeddings cluster turns by voice so each utterance is attributed, including overlapping speech on multi-party calls.
Captioning
A commit policy that keeps partials stable enough to read, with caption latency budgets set per channel and measured per session.
Vocabulary
Contextual biasing with domain lexicons — equipment IDs, drug names, tickers — and word-level confidence that routes uncertain terms to QA.
Italic grey = partial hypothesis, revised as audio arrives. Dotted underline + value = committed word with confidence below 0.85, flagged for the captioning QA queue.
Customer Lead Scoring
Scores get acted on when people trust them. Our lead models are transparent by construction: documented features, published weights and an exact attribution for every point, so a rep can see why a lead is Tier A and what would change it.
Lead score
72.4/ 100
Tier
B
Tier B action: Enrol in a technical SDR sequence
Feature attribution (points)
Score 72.4, tier B. Largest driver: Industry fit, +14.7 points.
How the attributions are exact. The model is logistic, so it is linear in log-odds: logit = β₀ + Σ wⱼ·fⱼ. For a linear model with independent features, the SHAP value of feature j is exactly φⱼ = wⱼ·(fⱼ − E[fⱼ]), and the φⱼ sum to logit − E[logit] with no approximation. To report points on the 0–100 scale, every φⱼ is multiplied by one shared factor — (score − baseline) ÷ (logit − baseline logit), the secant slope of the sigmoid — which preserves sign and ranking and makes the bars sum exactly to score − baseline. Rounding to 0.1 is apportioned by largest remainder so the displayed numbers add up too. This is a rescaling of log-odds SHAP, not a separate Shapley computation on the probability output.
| Feature | Transform fⱼ | wⱼ | E[fⱼ] |
|---|---|---|---|
| Employee count | log₁₀(employees) | 0.55 | 2.70 |
| Industry fit | fit score 0–1 | 1.60 | 0.50 |
| Annual revenue | band index 0–5 | 0.30 | 2.00 |
| Tech-stack match | match % ÷ 100 | 1.40 | 0.40 |
| Region | coverage fit 0–1 | 0.60 | 0.60 |
| Pricing-page visits | log₂(1 + visits) | 0.45 | 1.20 |
| Docs depth | log₂(1 + pages) | 0.25 | 1.50 |
| Demo requested | 1 if requested | 1.50 | 0.08 |
| Email engagement | click rate ÷ 100 | 1.20 | 0.18 |
| Days since last touch | days ÷ 30 | -0.55 | 1.50 |
| Intercept β₀ | — | -5.331 | — |
Tiers: A ≥ 75, B ≥ 55, C ≥ 35, D below. Weights are illustrative, fit to a synthetic population; a production model is refit on your closed-won history and monitored for drift like any other.
Features
Firmographic fit and behavioural intent from CRM, product analytics and marketing automation, with every transform (log, decay, cap) written down.
Calibration
Logistic or monotone gradient-boosted models, calibrated so a score of 70 corresponds to roughly 70% observed conversion on validation data.
Attribution
Exact SHAP values for linear models and TreeSHAP for boosted trees, written to the CRM beside the score so reps see why, not only what.
Monitoring
Score-distribution PSI and conversion by tier on control charts; the model is refit when calibration drifts, not on a calendar.
Two weeks, fixed scope. We map every model, prompt, data flow and decision path you run today, score them against your compliance regime, and hand you a ranked remediation plan with named mechanisms — not a slide deck.