The useful test of an AI coding workflow begins when the keynote instructions meet a real repository. If the demonstrated sequence fails there, buyers should treat the presentation as a capability claim under controlled conditions, not proof that the product fits their codebase.
On Tuesday, an engineering lead follows the keynote workflow verbatim. The prompts match. The sequence matches. The team’s repository does not compile.
That failure is more informative than a polished stage result. It exposes the distance between a workflow that can succeed in a prepared environment and one that can handle the dependencies, conventions, tooling and accumulated decisions inside production software.
A keynote demonstrates the chosen path
Product demonstrations answer a narrow question: can the vendor show this workflow succeeding under the conditions selected for the presentation?
That matters. A working demonstration carries more weight than a slide full of promises. But the vendor still controls the repository, task, environment and route through the product. The audience usually cannot see how many candidate tasks were rejected, how much setup happened beforehand or what the workflow does when its first assumption is wrong.
A real repository removes that control.
The code may depend on generated files, private packages, pinned toolchains, unusual build flags or tests that expect local services. Naming conventions may carry years of team history. A change that looks correct in one file can violate an interface somewhere else. Compilation catches only part of this, but it provides a useful first boundary: the proposed change cannot yet enter the normal engineering workflow.
The failure does not prove the tool has no value. It does invalidate the easiest interpretation of the demo, that another team can reproduce the result by copying the visible steps.
Compilation is the first checkpoint
Teams assessing coding tools often start by judging the generated code on screen. Does it look plausible? Did the tool touch the expected files? Is the explanation convincing?
Those checks can create false confidence. Code can look reasonable while importing the wrong module, calling an outdated interface or ignoring repository-specific build rules. The compiler has no interest in how persuasive the explanation sounds.
A stronger evaluation starts with the repository’s own gates:
- Run the documented setup process in a clean environment.
- Execute the workflow without hidden manual corrections.
- Build the affected targets.
- Run the relevant tests, linters and type checks.
- Inspect the diff for unrelated changes.
- Record every intervention needed to reach a passing state.
The intervention count deserves particular attention. If an experienced engineer quietly fixes three commands, supplies missing context and redirects the tool twice, the final output may pass while the advertised workflow still fails. Those corrections are part of the operating cost.
This resembles the problem described in The Friday the Checkpointing Code Becomes Optional: a reassuring success state means little when the control that makes it trustworthy can be skipped. For coding products, build and test gates must remain outside the tool’s self-assessment.
Measure recovery, not only first-pass success
A failed compile can reveal more than a successful one. The next question is whether the product can diagnose and recover from the failure without creating a wider mess.
Watch what happens after the error appears. Does the tool read the actual compiler output? Does it identify the relevant dependency or interface? Does it make a narrow correction, or start changing adjacent files until the error disappears? Can the engineer understand why each edit was made?
Recovery quality separates a useful coding assistant from a convincing code generator. Production work contains incomplete context, stale documentation and conflicting constraints. A product that performs well only when every assumption has been prepared upstream transfers the difficult work back to the team.
Keep the test bounded. Choose a task small enough for a human reviewer to understand, but connected enough to exercise real repository constraints. Preserve the initial prompt, generated diff, build output, follow-up prompts and final result. Then compare the tool-assisted attempt with the team’s normal process.
The goal is not to manufacture a failure. It is to discover where responsibility moves from the product to the engineer.
Put the real repository into the buying process
Before adopting an AI coding product, ask the vendor to support an evaluation on code that resembles your own constraints. If proprietary code cannot leave your environment, use an internal trial with an approved configuration and documented access limits.
Define success before starting. A useful scorecard can include first-pass build status, tests passed, engineer interventions, unrelated files changed, review time and whether the final patch is understandable. Avoid collapsing those observations into one vendor-generated score. The same concern appears in The Self-Grading Ad Platform Problem: the system making the claim should not be the only system judging the result.
Most importantly, preserve the failed attempt. A clean final diff can hide the route taken to produce it. The transcript and intermediate changes show how the product behaves when context is missing, an instruction conflicts with the repository or the first solution breaks the build.
The next evaluation should begin with the Tuesday failure still visible. Give the product the compiler output, limit the permitted files and count the corrections required. That record will tell the buying team more than replaying the keynote until it works.
Sources
No source links were supplied for the described workflow test.
Comments
No comments yet.