N° 017 Oct 20263 min read

AI does the work, never holds the verdict

Trusting without checking

Every day I see teams oscillating between lack of trust in AI and blindly following whatever the AI suggests, with very little space in between those two modes. The former are either change averse, grabbing on to their own working methodology, and what used to work in the past, while the latter are either lazy or excessively optimistic.

Focusing on this latter group, this leads to wild iterations and token spending, producing lots of code that may or may not be entirely needed or even correct, and making it more difficult both to validate and maintain.

Naysayers will argue that human coders are as much, or even more, prone to errors as AI, and I won’t disagree. However, the expectation is that those errors are eventually corrected, through careful iteration, debugging and simplification, not by brute forcing a solution by over coding.

At the same time, uncontrolled use of AI in SDLC moves the bottleneck from IT (analysis, design, development, integration testing) to acceptance, not addressing the somewhat inflated expectations of senior stakeholders who need their investments in this technology to start paying dividends.

Why the verdict matters

All of us have seen fantastic examples of applications and demos being “one shot” by the latest AIs, we’ve seen previously undiscovered vulnerabilities getting discovered and patched in record time, and a true explosion in terms of speed of development. However, at the end of the day, all of these are created based on a stated goal, and this goal is still human driven. And if the goal is human driven, the only true decider on whether it has been met must be the humans themselves.

We also have the accountability angle. Can an AI model be accountable for inaccurate, wrong, or unlawful actions? By today’s standards, of course not. Even without a “human in the loop”, someone made the decision of setting a goal, of letting a model loose, and accepting whatever action it takes with little to no control.

This is not acceptable. The outcome and the means matter. The verdict matters, and it still belongs to whoever is driving the AI.

One rule, and its cost

What does this mean for software delivery? It means focusing on what actually needs our input and critical thinking, and letting AI do what it does best. We’ve quickly adapted our ways of working and delegated too much critical thinking to AI. We’ve even used - and abused - AI to think creatively for us. While an interesting tool for getting your creative juices flowing, I think that’s misguided when used as an originator rather than an editor.

AI is stronger when we guide it properly, by providing goals and constraints and carefully examining its output, as it can sound very convincing even when defending “wrong” ideas. It’s also a very valuable partner as a sounding board, allowing for very interesting and in-depth sparring sessions that resemble a sort of accelerated design thinking session.

In software, this means using it to add detail to specs, turning those specs into code, creating test suites and automating those test suites. It doesn’t mean confirming that the specs are correct, or that the tests are exhaustive. Most of all, it can’t say whether the work that has been created actually serves the intended purpose.

This is a shift in the way we think about and source our teams. We need a stronger presence in the initial conceptual phases and, most of all, in the validation of outputs, which means (again!) changing the way we work, what activities are actually most valuable, and what to look for when building a high-performing team and company.

So one rule: critical thinking and decision making belongs to humans. In a sentence: use - and abuse - AI, but own the output.

← All notes