3 Comments
User's avatar
Felipe A. Zubia's avatar

I like this framing, I haven’t seen it put quite this way, and I think it’s a helpful perspective. That said, we can’t fully absolve the model itself of responsibility.

We accept fallibility everywhere, but always relative to risk e.g. duct tape on a leaky faucet is ok but it isn’t tolerated in a spacecraft. AI may not be the spacecraft (yet), but its well above the “faucet repair” tier, which means there’s a limit to how much agents, workflows, and post-hoc constraints can compensate.

At some point, duct tape stops working, and reliability has to be addressed at the foundational architecture and training level, where the problem truly resides.

Wren Dougherty's avatar

As much as I refuse to accept there’s any problem duct tape can’t solve 😄 - this is a great point.

The level of rigor, oversight, and system design absolutely depends on the use case, and there are plenty of contexts where “duct tape” workflows just won’t cut it. At some point, you need the underlying model improvements too.

I actually went back and rebuilt a few repos from my Feb-April Claude experiments using the current 4.5 models, and… wow. The difference is night and day. The surrounding systems matter a lot, but the base model is still doing the bulk of the work in determining whether the whole thing succeeds.

Structure helps you capitalize on better models, but certainly doesn’t magically replace them, or let underpowered models tackle many tiers of problems.

Sean's avatar

Thanks for sharing these insights Wren. I’ve been evaluating different techniques to push my company’s usage of AI tools further and I really like the way you’ve summarized these principles and capabilities. We’re still in the copilot mode and are wanting to evolve more to the coordinator mode of operating. I’ll share this with my team!