A programmer working on code with a laptop and monitor setup in an office.

Photo by Jakub Zerdzicki on Pexels

An AI feature has a moat when its value survives changes to the prompt, model provider, and surrounding interface. If swapping those parts removes most of the value, the feature depends on rented capability rather than a defensible product advantage.

A useful audit starts with subtraction. Remove the polished demo, replace the model, expose the workflow underneath, and measure what remains. The answer may be proprietary data, accumulated user feedback, workflow ownership, distribution, or trust earned through reliable performance. Sometimes little remains. Finding that out before a competitor does is the point.

Start with the replaceable layer

Write down every component required to produce the feature’s result:

  • The model and API
  • The system prompt
  • Retrieval sources
  • User-provided context
  • Product integrations
  • Evaluation rules
  • Human review
  • Stored feedback and corrections
  • The interface where users act on the result

Now mark anything a capable team could reproduce from public documentation in two weeks. Be strict. A prompt can be long, carefully tuned, and commercially useful while remaining easy to copy. The same applies to a chat interface, a standard retrieval pipeline, or a workflow that sends text to a model and displays the response.

This does not make the feature worthless. Replaceable features can acquire users, support a larger product, or validate demand. The mistake is treating current usefulness as evidence of durable protection.

Next, swap the prompt. Give another team member the intended input and output, but hide the original instructions. Ask them to recreate the behavior using the same model. If they reach an acceptable result within a day, prompt design probably contributes execution quality rather than a moat.

Then swap the API. Use a comparable model from another provider without rebuilding the rest of the product. Record changes in accuracy, latency, cost, failure modes, and user preference. The Friday Model-Swap Triage offers a practical way to treat that change as an operational test instead of a vendor debate.

Measure what survives the swap

A defensible feature should preserve meaningful value after its most replaceable components change. Test that claim with the same set of real tasks before and after each swap.

Choose 25 to 50 examples drawn from actual use, including failures, awkward edge cases, and inputs that require product-specific context. A showcase set will tell you whether the demo still looks good. A representative set will tell you whether the product still works.

Score outcomes against criteria users care about. Depending on the feature, that might include factual accuracy, edits required, task completion, response time, or the percentage of outputs accepted without correction. Avoid asking the model to grade itself. That can hide systematic errors behind confident scores, a problem examined in The Self-Grading Ad Platform Problem.

Run four versions:

  1. The current product with its current model and prompt.
  2. The same product with a replacement model.
  3. The replacement model with a newly written prompt.
  4. A bare alternative that receives only the user’s input and generally available context.

The gaps matter more than the absolute scores. If version four performs almost as well as version one, the underlying model supplies most of the value. If versions two and three remain strong because your data, integrations, evaluation system, and feedback history carry across, you have evidence of a sturdier position.

Look for compounding assets

The strongest protection usually sits outside the model call.

Product-specific data can matter when competitors cannot easily buy, scrape, or regenerate it. Raw volume alone proves little. The useful asset is often structured context tied to outcomes: which recommendation a user accepted, what they changed, whether the task succeeded, and why an output failed.

Workflow ownership can also persist through a model swap. A feature embedded in approvals, permissions, records, and downstream actions solves more than text generation. Replacing it means rebuilding operational connections and retraining users, not copying a prompt.

Evaluation systems are another durable layer. Teams that can detect regressions across hundreds of representative cases can adopt better models faster and reject attractive failures sooner. That capability grows as real usage produces more edge cases and corrections.

Distribution and trust count, with caveats. Existing access to the right buyers can lower acquisition costs. Reliable handling of sensitive inputs can influence purchasing decisions. Neither excuses a weak product, and neither should be claimed without evidence.

Use a simple test for each proposed moat: does it improve as more customers use the feature? If the answer depends on collecting data, confirm that customers have consented to that use and that the data produces measurable improvement. A vague promise of a future data flywheel is not current protection.

Turn the audit into a shipping decision

Classify the result honestly.

A wrapper is a feature whose core result can be recreated by combining a public model with a competent prompt and basic interface. A productized wrapper adds useful workflow, support, reliability, or distribution, but remains vulnerable to model providers and nearby competitors. A compounding product improves through assets that persist across model changes, such as proprietary outcome data, embedded workflows, trusted controls, or a tested evaluation corpus.

All three can generate revenue. They call for different decisions.

A wrapper may deserve a fast launch and limited engineering investment. A productized wrapper needs strong positioning, efficient distribution, and close tracking of platform changes. A compounding product merits deeper investment because each customer, correction, and integration can make the next replacement attempt harder.

Schedule the audit before the next roadmap review. Give one engineer two days to replace the prompt and provider, then rerun the evaluation set. Put the score changes, migration time, and surviving assets on one page.

If the feature collapses, you have found the dependency while there is still time to fix it. If it holds, you can name the reason with evidence.

Comments

No comments yet.