CTOs, heads of engineering and product leaders
Putting AI to work in your engineering and your product
In short
AI is two different jobs. One is helping your engineers use it well, without drowning in code nobody can review. The other is building AI features customers can rely on, at a cost that makes sense. I do both, and contribute to Deepnote’s open-source notebook, which has extensive agentic features.
You might recognise this
- Your engineers write far more code with AI, and review and CI cannot keep up.
- Coding agents ignore your team’s conventions.
- Your AI features default to the most expensive model, and the costs are climbing.
- You cannot tell whether an AI feature is good enough to ship.
Two different jobs
- Helping engineers use AI well
- AI code review with tools such as CodeRabbit and Codex, AGENTS.md files that teach coding agents your conventions, a shared channel for what works, and CI/CD scaled for the volume of code AI now produces.
- Building dependable AI products
- Evaluations that show whether a feature is good enough, or thorough manual QA with the risks stated plainly where tasks are too varied for them; models chosen for the task rather than defaulting to the most expensive, including purpose-built decision models for classification and routing; and oversight so customers can rely on the result.
Both are measured. For engineering: review load, CI runtime and change failure rate. For products: evaluation results and cost per request.
How we might work together
Examples, not packages. They often overlap within one relationship, and we shape the work together.
- AI-assisted code review
- Introducing automated review with CodeRabbit and Codex, so people review what matters.
- Shared conventions
- Maintaining AGENTS.md files across repositories, and a channel where engineers share what works.
- CI/CD for more code
- Scaling pipelines for the volume AI produces, and improving their runtime continually.
- Cost-effective AI features
- Evaluations first, then the least expensive model that passes them, with the most capable models kept for where they earn their cost. Classification and routing go to purpose-built decision models such as TypeSafe’s Jev.
Relevant experience
- More than 100 merged pull requests since September 2025 to deepnote/deepnote, Deepnote’s open-source notebook with extensive agentic features
- Helping move Deepnote from a data science platform to an agentic AI workspace
- Introduced AI code review with CodeRabbit and Codex, shared AGENTS.md conventions across repositories and scaled CI/CD for AI-assisted development
- Coding almost every day, running work in parallel on remote servers and my own machine, and asking engineers each week what they are doing differently
How we start
We start with a focused look at what you’re trying to achieve, what’s getting in the way, and how I could help. From there, we agree whether there’s a useful role for me and what it should involve.
- Understanding the business and what you are trying to achieve
- Meeting the people who matter to the work
- Looking at enough evidence to see where I could be useful
- Deciding together whether an ongoing role makes sense, and what it should be
Selected results
| Organisation | What changed |
|---|---|
| DeepnoteFractional CTO. YC-backed, $23.5M raised | Moving the product from a data science platform to an agentic AI workspace, raising AI-native engineering standards and improving enterprise delivery. Server costs cut by 40%. |
Questions
Which AI coding tools do you recommend?
Whichever measurably helps your team. I have introduced CodeRabbit and Codex for review; the choice follows the evidence, not the hype.
Do we need the most capable model?
Rarely for everything. Evaluations usually show that a smaller, less expensive model handles most requests, with the most capable models kept for where they make a measurable difference.