Skip to content

Tools Constellation

Every tool we run, and exactly what for.

22 models, agents, media generators, clouds and design tools — 12 in production on engagements, the rest under evaluation or in sandboxed R&D. Each entry says what we use it for, not what the vendor says it does.

Index

The full tools index

Filter by category or maturity, or search by what a tool is used for. Every card has a stable anchor, so /tools#tool-claude links straight to it.

Category

Maturity

Showing 22 of 22 tools

  • Claude

    Production

    Foundation Models · Coding · Agents · Data

    Our primary model for agentic coding, spec drafting, long-context document review and structured extraction, called through private endpoints with every prompt hashed to the ledger.

    Product page for Claude (opens in a new tab)
  • Claude Plugins

    Production

    Agents · Coding

    We package our review commands, hooks and connectors as versioned plugins so every engineer's coding agent runs the same governed toolchain.

  • Claude Skills

    Production

    Agents · Coding

    We encode house procedures — eval runs, landing-zone reviews, ledger export — as skills the agent loads only when the task calls for them, keeping context budgets tight.

  • Cursor Rules

    Production

    Coding

    Repository-scoped rule files, reviewed like code, that pin architecture boundaries, banned APIs and secret-handling patterns into every Cursor session.

  • Jev

    Evaluating

    Agents · Coding

    Agent tooling we are running through our coding-agent harness on internal repositories to decide whether it earns a place in the build pipeline.

  • Hermes 2.0

    R&D

    Foundation Models

    An open-weight model line we benchmark in our eval harness for scenarios where inference has to run self-hosted, inside a boundary no external API can reach.

  • OpenClaw

    R&D

    Agents

    An open-source agent runtime we study in an isolated sandbox to understand tool-permission and prompt-injection behaviour before anything similar touches client systems.

  • Seedance

    Production

    Video/Image

    Video model we used, via Higgsfield, to generate this site's hero backdrop; we use it for motion studies and explainer footage.

  • Kling

    Production

    Video/Image

    Video model we used alongside Seedance, via Higgsfield, to generate this site's hero backdrop and ambient motion loops.

  • Nano Banana

    Evaluating

    Video/Image

    Image model wired into our media pipeline as the keyframe generator; it is not yet enabled on our API account, so keyframes route through other models today.

  • Grokbot

    R&D

    Agents

    Conversational agent tooling we explore in a sandbox to compare agent behaviour and guardrail responses across model providers.

  • Base 44

    Evaluating

    Coding · Data

    An app builder we are evaluating for internal tools that a client's business team could own and maintain after hand-off.

  • Auri.build

    R&D

    Coding

    An AI build tool we are assessing against our spec-driven workflow; it has no production role yet.

  • AI Figma

    Production

    Design

    Figma's built-in AI features, which we use to explore layout variants and draft prototype flows before a design review — the reviewed file, not the suggestion, is the source of truth.

    Product page for AI Figma (opens in a new tab)

Evaluation rubric

How we evaluate tools

A tool moves from R&D to Evaluating to Production only by clearing the same five gates. Nothing is grandfathered: model and vendor updates re-run the rubric.

01 · Security review

Threat model before first prompt

We map where prompts, files and credentials go, check SSO and SCIM support, audit-log export and token scoping, and red-team the tool with our prompt-injection suite before it sees anything but synthetic data.

02 · Data residency

Know which region holds the bytes

We confirm processing and storage regions, retention and training-use terms in writing, and prefer deployments inside the client's own tenancy reached over private endpoints.

03 · Eval-harness scores

Scored on our golden sets, not demos

Each candidate runs the same versioned task suites — coding, extraction, reasoning, refusal behaviour — and must beat or match the incumbent within a declared tolerance band on every suite.

04 · Cost

Unit economics per task, not per seat

We measure tokens, latency and retries on the eval suites and project cost per completed task at the client's volume, including the human review time the tool's error rate implies.

05 · Exit path

Replaceable by design

Every tool sits behind our orchestration layer's provider interface, and prompts, rules and eval sets live in the client's repository — so swapping a vendor is a config change plus an eval run.

Production
Passed security review and our eval harness; approved for client engagements inside the client's own tenancy, with outputs logged to the decision ledger.
Evaluating
Running against our golden sets and cost model on synthetic or internal data. Not used with client data until it clears the rubric.
R&D
Studied in an isolated sandbox to understand behaviour, failure modes and security posture. No client data, no production path.

Choosing between models and vendors?

We run your candidate tools through the same rubric — security review, residency, golden-set scores, unit cost and exit path — and hand you the evidence, not a vendor ranking.