What over a year of building with AI has taught me

Claude Certified Architect Stamp

I recently earned the CCAR-F certification, which gave me a good reason to stop and take stock of the past year. Most of that year has been spent learning how to build with AI tools rather than just talking about them, and the way I work has changed quite a bit in that time. I thought it was worth writing down what I’ve actually been using and the handful of habits that have made the biggest difference.

How the certification came about

The CCAR-F came through an opportunity at Entelect, and I’m grateful for it. This was (is?) not available to everyone, so the opportunity to take it and learn from it has been super valuable. Sitting it also forced me to spend deliberate time on the material instead of picking things up as I went, which is a different kind of learning.

The certification covers two areas that map closely to each other: building applications with AI, and designing applications that use AI as part of how they work. I should be honest and say that I’m not an expert in either after one certification. What it did give me was a clearer vocabulary for the problems and a bit more confidence about what I don’t know.

Where my AI use has gone

Those two jobs are very different, and so far I’ve mostly done one of them. Building software with AI is where almost all of my experience sits. Building software that uses AI is something I’ve only done a small amount of, and it’s probably the thing I’m most interested in exploring next.

On the personal side, my use is pretty simple. Gemini through the web is still my default for research tasks - summarising something I don’t want to read in full, comparing options, or just getting a quick answer without the noise of a search results page. Antigravity has become my main tool for building personal projects and day-to-day tasks, mainly due to the generous limits they offer as part of subscriptions I already have!

The majority of my use, though, is in software engineering. Antigravity, Claude Code, opencode and pi are all in rotation, and which one I pick usually comes down to the task in front of me or what the environment allows, rather than loyalty to any of them. More recently I’ve been spending time with agentic harnesses and assistants like OpenHands and Hermes, mostly to see where the boundaries are and what changes when the agent is running more of the loop on its own.

Building software with AI and building software that uses AI are different skills, and I’ve spent far more time on the first one than the second.

Tips and tricks I’ve picked up

Code is compiling XKCD
Click to open in new tab - XKCD #303

These are the things I’d tell someone starting out, and the things I keep having to remind myself of.

Work in small slices, and still review

Almost every lesson I’ve learned this year comes back to this one. Big sweeping prompts feel productive and rarely land well, particularly when you need very specific, rule oriented outcomes. I get much better results when I break work into small, self-contained slices and direct the tool one slice at a time, checking as I go.

And I still review what comes out. The models are frequently good, but I don’t think they’re at the point where I’m comfortable handing over the keys, and I’m not sure I want them to be. I care about the implementation details, and the review step is where I learn the codebase and catch the things that would have caused problems later. There’s a practical side to it too - I’m the one who has to maintain this, so I want to understand it.

Experiment with the models

As you’d expect, the more capable models when it comes to reasoning tend to be better at planning out large chunks of work. If I need a plan or a design, I’ll reach for a stronger model.

For the implementation that follows, I usually go the other way. I’ll often pick a faster, cheaper, “dumber” model instead. The reason is the feedback loop. I’d rather have a quick back-and-forth where I steer it back on track, even if that takes a few more turns, than wait on a slower model that gets there in fewer steps. There’s still back and forth either way, and I’d rather that process be as fast as possible. A model that’s slightly better on paper but dramatically slower rarely wins that trade for me. Concretely, I tend to favour something like Gemini 3.8 Flash or DeepSeek V4.1 Flash for implementation over Sonnet 5 or GPT 5.6 Terra. The output is good enough, and it’s dramatically faster.

Tip

Match the model to the job. Planning wants reasoning depth. Implementation usually wants speed, because you’re going to review it anyway.

CI/CD and automated tests matter more than ever

This one surprised me the most. A solid pipeline with automated builds, unit tests, integration tests and end-to-end tests is valuable for all the usual reasons - it keeps the system stable and correct, and it catches regressions before your users do. What’s changed is that it’s now also the feedback mechanism for the agent. When the tests run quickly and reliably, an agent can run them, see what broke, and fix it on its own. The better your safety net, the more of the loop the agent can close.

Architecture tests have been a nice find here too. They’re a cheap way to keep a codebase from drifting into a structure nobody agreed to, and they give the agent a clear signal when a change would break a boundary. If you haven’t looked at them before, they’re worth an afternoon.

Tip

Speed is part of the feedback loop. A test suite that takes twenty minutes will stall every agent iteration that depends on it, so its speed is worth protecting as much as the build itself.

Where I’ve landed

Note

AI is an accelerator. If you point it at a wall, expect to hit it very fast (with consequence). Your role is to direct it to the safe path.

A year ago I was still skeptical about how much of this would stick. I’m not anymore, but my position is more specific than “the tools are great”. The tools improved, but most of the difference came from the habits around them - reviewing everything, working in small slices, and investing in the tests and pipeline that let me move with confidence. That last one pays off no matter which tool or model is winning next quarter. Worth noting - the benefit is real, you can move a lot faster with generative AI tools, as long as your guard rails are in place.

If you’ve been experimenting with any of this yourself, I’d love to hear what’s worked for you, especially if you’ve spent more time on apps that use AI at runtime than I have. That’s the part I’m most keen to dig into next!