AI enablement · Strategy through delivery

Everyone has an AI demo. Almost no one has an AI system.

The distance between those two things is where AI projects go to die — and it has almost nothing to do with the model. We close that distance. We write the code, and we sit in the room where you decide whether the code is worth writing.

Diagnostic · 2 weeks Production · 6–10 weeks Evaluation set before code Your team owns it
01The gap

The model was never the hard part.

Capability is now the cheapest input in the stack. You can rent a model that reasons better than most of your process documentation, and you can have it doing something impressive by Thursday. What you cannot rent is the unglamorous work between that demo and a system the business can actually stand on.

State A

Works in a demo

  • One curated file
  • One happy path
  • One user, and it is you
  • No permission model
  • No cost ceiling
  • Judged by whether it feels impressive
The part nobody scoped

Where the eight months go

  • Data plumbing across systems
  • An evaluation set with real answers
  • Retrieval that respects authorization
  • Failure detection and fallback
  • Cost and latency budgets
  • The workflow redesign underneath
  • Rollout, training and an owner
State B

Works on Monday

  • Forty edge cases you did not invent
  • Real permissions, per user
  • People who never asked for this
  • An invoice with a comma in it
  • A number the CFO can check
  • Judged by whether anyone still uses it

Nothing in the middle column is exotic. All of it is knowable in advance. Almost none of it appears in the plan, because the plan was written by people who had only ever seen state A.

02Failure modes

Seven ways it breaks. None of them are the model.

This is the list we walk through on a first call. If three or more are unanswered, you do not have a model problem — you have a scoping problem, and it is enormously cheaper to find that out now than in month four.

01

The data is not ready

It lives in six systems, two of them have no usable API, and the one that actually matters is a shared drive of scanned PDFs with a naming convention from 2014. Every AI timeline that skipped this step is now a data-engineering timeline wearing a disguise.

02

Nobody defined correct

Without an evaluation set, “it seems better” is your only quality signal. You cannot tune, defend, or safely change a system you can only judge by feel, and every prompt edit becomes a coin flip nobody is allowed to lose.

03

Ninety-two percent is not a product

A model that is right nine times in ten makes a spectacular demo and an expensive support queue. Essentially all of the product work lives in the tenth case: detecting it, containing it, and routing it to a human before it reaches a customer.

04

Permissions are the actual project

The moment retrieval can read everything, the assistant can say anything to anyone. An index without an authorization model is a data breach with a friendly interface, and retrofitting per-user access into a working prototype is usually a rewrite.

05

The bill scales with usage, not seats

Tokens, times real volume, times retries, times all the context you quietly re-send on every call. Almost no pilot runs that arithmetic, so the number arrives fully formed in month three and the conversation stops being about capability.

06

The workflow never changed

Bolting a model onto a broken process automates the broken process, faster and at higher cost. A large share of the value in any credible AI project is a process redesign wearing an AI costume, and that part is organisational work, not engineering.

07

It works for the person who built it

Adoption is not a launch email and a wiki page. If the new way is not visibly faster than the old way on the first attempt, on a normal day, under normal pressure, your team is quietly back in the spreadsheet by Tuesday.

03Why us

Most people selling this can only do half the job.

The market splits cleanly in two, and both halves fail in a way you can predict from the org chart.

Half one

The strategy side

Firms that can name the opportunity and have never shipped a system. You get a maturity model, a roadmap, a governance framework, and a slide that says leverage AI across the value chain. The estimates are guesses, because nobody in the room has ever had to honour one.

Result: a plan nothing can be built from
Half two

The build side

Shops that can wire an API in a fortnight and have never carried a number. You get a working prototype for a workflow that was not worth automating, no evaluation harness, no cost model, and no honest answer to whether the thing is actually working.

Result: a system nobody uses

The job is the seam between them.

The people who write the retrieval layer are in the meeting about whether the retrieval layer should exist. Same week, same team, no translation loss, no handover between whoever scoped it and whoever has to make it true.

What the business said What it actually means What gets built
SaidWe need to cut support costs.
MeansDeflect the third of tickets that are already answered by twelve help articles — without deflecting the ones that churn a customer.
Builtretrieval over 40k resolved tickets
per-tenant ACL filter at query time
confidence gate → human queue
eval set: 250 tagged tickets
SaidWe need AI in the product.
MeansFind the one place a model is genuinely better than the form we already ship, and resist the other eleven.
Builtdraft generation inside the existing editor
streamed, editable, never auto-saved
accept-rate tracked per user
kill switch per account
SaidOnboarding takes too long.
MeansFour of the ninety minutes are manual transcription from documents. The other eighty-six are a process problem no model will touch.
Builtstructured extraction from uploads
schema-validated, typed output
low-confidence → review queue
and a memo about the other 86 minutes

Enclave is deliberately small, and that is the point.

You get the people who write the code in the meeting where the decision is made. No account manager relaying between them. No bench of juniors learning the failure modes on your budget. No gap between what was sold and what can be delivered, because the same people are on the hook for both.

We have spent our careers on both sides of this — building the systems, and being accountable for whether they were worth building at all. That is an unusual combination to find in one room, and it happens to be exactly the combination this problem needs.

01

We will tell you when AI is the wrong tool.

02

Nothing ships without an evaluation set.

03

Scope stays narrow until something is live.

04

Your team owns it when we leave.

04Engagements

Three ways in. Each ends with something you can act on.

Fixed scope, fixed duration, and a stated deliverable — so the thing you are buying is legible before you buy it.

2 weeks · fixed fee

Diagnostic

Where AI actually pays here, and where it will quietly waste a year.

  • Every candidate workflow scored on value and on feasibility, separately
  • Data readiness assessed per workflow, not in the abstract
  • A build, buy, or don't call on each one, with reasoning attached
  • One sequenced recommendation with a cost and a timeline you can take to a board
Fit

You have budget and pressure and no defensible way to choose what to do first.

Not

You already know the workflow and just need it built. Start at the next one.

6–10 weeks · fixed scopeCore

Pilot to production

One workflow, taken all the way to a normal Monday morning.

  • An evaluation set built from your real cases, with a bar agreed before any building
  • The system itself: retrieval, orchestration, permissions, fallbacks, the unglamorous parts
  • Instrumentation for quality, cost and latency that your team can read without us
  • A staged rollout, a runbook, and the handover that makes it yours
Fit

A pilot that stalled, or a workflow you already know is worth the money.

Not

Eight workflows at once. Pick one. The second is always cheaper than the first.

Monthly · capped hours

Embedded

For a team already building, and drifting slightly sideways.

  • Architecture and vendor decisions made with people who have made them before
  • Evaluation discipline installed in your CI, not written down in a document
  • Review on the AI surface area: prompts, retrieval, tool boundaries, failure paths
  • The unpopular no, delivered early enough that it still saves you something
Fit

Strong in-house engineers, and nobody senior who has shipped this particular thing.

Not

You want someone to write all of it. That is the engagement to the left.

05Method

The order matters more than the tooling.

Five steps, run in this sequence, on every engagement. The sequence is the method — most failed projects ran the same five steps in the wrong order.

01

Map

The workflow as it actually happens, not the version in the process document. Where the time goes, who touches it, and what a better outcome is worth in money.

02

Define correct

Build the evaluation set before the system. A few hundred real cases with the right answer attached. Everything after this becomes measurable instead of arguable.

03

Thin slice

The narrowest end-to-end path that touches real data and real users, in production behind a flag. Not a notebook, not a sandbox, not a deck.

04

Instrument

Quality, cost, latency and fallback rate on a dashboard your team reads weekly. A system you cannot see is a system you do not own.

05

Hand over

Documentation, runbook, the evaluation suite, and enough context to change it safely. The engagement succeeds when you stop needing us.

Step two is the one everybody skips, because it feels like a delay and produces nothing demo-able. It is also the only reason step four is possible, and the difference between a system you can improve and one you can only argue about.

06Capability

Chosen per workload. Never per press release.

No vendor relationships to protect and no house framework to sell you. The architecture is argued from your constraints, and it is written down with the reasoning intact so your team can revisit it later.

Models

Frontier, open-weight, or local

Frontier APIs where reasoning quality decides the outcome, smaller models where volume and latency dominate, open weights where the data cannot leave. Decided against your evaluation set, not a benchmark table.

Retrieval

Search that respects who is asking

Hybrid retrieval, chunking that respects document structure, and per-user authorization applied at query time rather than filtered afterwards. This is where most systems are quietly broken.

Evaluation

Golden sets and regression gates

Human-labelled cases, model-graded scoring calibrated against those labels, and a regression suite that runs in CI so a prompt change cannot silently cost you six points of accuracy.

Orchestration

Tools, agents, and knowing when not to

Most problems presented as agentic are a well-specified pipeline with two branches. Agents earn their unpredictability only when the path genuinely cannot be known in advance.

Deployment

Your cloud, your VPC, your hardware

Including fully on-premise and air-gapped, with open-weight models and an audit trail that never leaves your perimeter. A long-standing part of this practice, and a constraint best established before the architecture.

Safety & cost

Guardrails, budgets, and blast radius

Handling of sensitive fields, prompt-injection surface where a model reads untrusted text, spend ceilings per tenant, and a fallback that degrades to something useful rather than to an apology.

07Questions

Before you ask.

Judgement they have not had the chance to build yet. Good engineers get an AI system to eighty percent quickly, then spend months on the rest, because the failure modes are unfamiliar: evaluation, retrieval quality, permission boundaries, cost at real volume, and knowing which of those to fix first. We have made those calls before, and we make them alongside your team rather than around it.
A consultancy sells you a strategy and hands the build to somebody else. An agency sells you a build and never questions the strategy. Enclave does both, so the estimate comes from the people who have to honour it, and the architecture comes from people who understand what the business is actually paying for.
Both, usually in the same week. The diagnostic is analysis and produces a decision. The production engagement is hands on the keyboard: the retrieval layer, the evaluation harness, the fallbacks, the instrumentation. Your team ends up owning all of it, which is the point of the last two weeks.
You will hear that early, often on the first call, and it costs you nothing. A surprising share of what gets scoped as an AI project is a reporting problem, a data-quality problem, or a process problem that a model will make faster and no better. Saying so is cheaper for both of us than discovering it in month four.
Whichever the workload justifies. Frontier APIs where reasoning quality decides the outcome, smaller or open-weight models where volume and latency dominate, and fully local deployment where the data cannot leave. The choice is argued against your evaluation set and your cost ceiling, and it is written down so it can be revisited when the market moves again.
Yes, and it is a substantial part of this practice. Systems can run inside your cloud tenancy, your VPC, or on hardware you own, including fully air-gapped, using open-weight models with an audit trail that stays in your perimeter. The engineering is different rather than impossible, and the constraint belongs in the conversation before the architecture, not after it.
The diagnostic is a fixed fee for a fixed two weeks. Production work is scoped and quoted before it starts, so the number does not move unless the scope does. The first call costs nothing and is the fastest way to find out whether any of this is worth doing.
A thin end-to-end slice running against your real data, behind a feature flag, is normally live within the first few weeks of a production engagement. That is deliberate. A narrow slice in production tells you more in a week than a broad prototype tells you in a quarter, and it tells you the expensive things first.
→Start here

Start with a call, not a proposal.

Thirty minutes. You describe the workflow. We tell you whether it is viable, what it would genuinely take, and whether you need us at all. If the honest answer is that AI is the wrong tool here, you get that answer on the call.

Thirty minutes, no deck, no sales engineer You leave with a viability read either way Under NDA on request, before any detail Reply within one business day
Please enter your name.
Enter a valid work email.
Please enter your company.
Please choose one.

You are agreeing only to be contacted about scheduling. Nothing you write here is shared, and an NDA can be signed before any detail is exchanged.

Received.

You will hear back within one business day with a couple of times. If you included the workflow, the first call will already be about that rather than about introductions.