The work is still too vague
A promising demo does not tell you who owns the process, which exceptions matter, or what a useful result would change.
Gary Butler · AI product engineering
I help small teams decide whether AI belongs in a real workflow, then build the smallest version that can prove it. You get something your team can run, inspect, and make a decision about.

A model demo can look convincing in an afternoon. These are the questions that show up when a team tries to turn it into useful software.
A promising demo does not tell you who owns the process, which exceptions matter, or what a useful result would change.
Real work arrives through documents, systems, and people. The proof needs safe examples and a clear source of truth.
The team needs to know what the model may suggest, what stays deterministic, and where a person must approve the result.
A useful proof leaves behind code, test results, operating cost, and enough context for the team to make the next decision.
Bring one repeated process that costs time, creates rework, or depends on one person remembering how everything fits together. We agree on the definition of done before work starts.
See the proof sprint →We document who does the work, what information they use, where it slows down, and which result matters.
I build the smallest version that can test the riskiest part of the idea against representative examples.
You see where the proof works, where it fails, how much each run costs, and which decisions still need a person.
The handoff says whether to build, buy, narrow the problem, or stop. You also receive the source code and setup notes for anything I build.
These are experiments, not client case studies. Each one records the question, the boundary, what exists now, and what still needs proof.
Visit the lab →Can an agent turn a noisy set of sources into a short brief without hiding where its claims came from?
Can a model help shape an early product idea without inventing certainty?
Can model selection be explained by the job, cost, latency, and failure risk instead of a leaderboard?
Practical thoughts on prototypes, agent boundaries, evaluations, and the decisions that matter after the demo.
View all writing →Autonomy gets the attention. Good judgment—and a clear boundary—earns the trust.
Read essay ↗A faster way to learn whether an AI idea is valuable before building the machinery around it.
Read essay ↗The leaders making the best AI decisions are building a habit, not memorizing a vocabulary.
Read essay ↗I have spent more than 30 years building software and 17 years leading development and analyst teams. That work taught me to care about permissions, failure paths, handoffs, and the people who support the system later.
View my experience →My background includes enterprise applications, integrations, modernization, and production support. I work with the systems a team already owns instead of pretending the old world does not exist.
My recent work covers LLM applications, retrieval, tool-using agents, structured outputs, evaluations, and human approval. I use those tools when they fit the job.
Start with the work
Send me a short description of the workflow, who owns it, and what keeps going wrong. I will tell you whether a proof sprint makes sense and what I would test first.
Describe the workflow ↗