AI Across the Engineering SDLC – Case Study

AI Across the Engineering SDLC -
A Cleverbit Client Case Study

Building AI into engineering that’s consistent, traceable, and safe to rely on… and measuring what it delivered.

AI Across the SDLC

CASE STUDY

Like most engineering teams, this client had started using AI in the way that felt most natural: individuals prompting a chat tool when it was useful. The results varied by person and by prompt. It helped in places, but it didn’t move the consistency of what the team shipped, and it introduced a new failure mode: AI-generated code that looked right but wasn’t grounded in how the product actually worked.

The client is a strong engineering organisation, and it wanted to use AI deliberately rather than incidentally.

The goal was to build AI into the way the team already works: consistently, traceably, and safely enough to rely on.

Cleverbit worked with the client’s engineering organisation to design and build that capability. 

"Choosing the right model early was the unlock: once sub-agents and skills were in place, it stopped mattering which engineer wrote a given piece. The output held together in a way it hadn't before."

In this article

How we approached the situation

Three decisions shaped the whole engagement.

      1. Choosing the right foundation. The client worked with Cleverbit’s R&D team to evaluate available models against their actual work, running tests focused on their own context, and backed the provider that held up better in practice. That gave the team a head start as sub-agents and skills became viable.
      2. Keeping humans in control at every gate. AI drafts, proposes, and reviews. Engineers and QA approve. Every skill that creates an artefact has explicit sign-off checkpoints; review and analysis skills are read-only. Establishing this as a cultural norm, not just a technical constraint, was a deliberate part of the work.
      3. Keeping humans in control at every gate. AI drafts, proposes, and reviews. Engineers and QA approve. Every skill that creates an artefact has explicit sign-off checkpoints; review and analysis skills are read-only. Establishing this as a cultural norm, not just a technical constraint, was a deliberate part of the work.

"One well-scoped skill did most of the heavy lifting on the repetitive changes, enough context to get each one most of the way, leaving engineers the last stretch rather than the grind."

What we built

Over the course of the engagement, ten uses of AI were embedded across the client’s SDLC, each built as a version-controlled skill, a sub-agent workflow, or a custom application on top of the connected context.

Requirements and component design

AI now drafts requirements documents and component designs against consistent writing and quality rules. It has already caught duplicated and misassigned work, flagging when a component overlapped another team’s area before it was built twice.

Large-scale code transformation

The team regularly needs to apply the same structural change consistently across a large codebase. A skill was built with just enough context: the rationale, the ways of working, and the necessary guardrails. It gets each change most of the way, then a developer takes it to done. The consistency gain matters as much as the time saving.

Code consistency across teams

The client’s engineering teams came from different backgrounds, each with their own conventions. Shared conventions are now encoded once in the skill library and available to everyone, so the same standards apply regardless of who is writing the code.

Agentic code review

Peer-review skills tuned to the client’s stack now check for regressions, performance issues, null-safety, security, and design quality, not just whether a piece of work meets its acceptance criteria. This takes load off other developers and catches issues before they reach merge.

Test strategy

The AI-assisted test strategy draws on acceptance criteria, past incidents, historical defects, and review threads. It questions the design rather than just confirming it. A QA engineer still signs off before anything is created; the agent runs a multi-step chain from strategy through to test creation with mandatory human gates.
Agentic SDLC
 

Onboarding and local environment setup

New developers now get AI-assisted setup grounded in the repository’s own configuration and documentation. The long tail of environment-specific issues that previously consumed senior time has largely been absorbed.

Automated release notes

Detailed, customer-facing release notes are generated automatically from one source of truth, kept consistent and complete without manual assembly.

In-house troubleshooting assistant

Engineers and support staff ask questions in plain language and get back likely causes, troubleshooting steps, known workarounds, the linked defect, and the release that fixed it. The search covers the client’s entire support and engineering history, including the discussion and resolution threads where the real answer usually sits. Every response cites its sources. The client is exploring how to extend it beyond a single team.

Live production root-cause analysis

An AI agent pulls live telemetry alongside the source code and matches what production is doing to the code paths that could explain it. Applied to a production incident, it narrows the problem to a small number of evidence-backed hypotheses, identifies the candidate configuration and code for each, and proposes fixes for engineers to evaluate. It keeps proven findings separate from hypotheses, so the output is a reliable starting point rather than a guess.

Cloud cost analysis

The client measured savings from an infrastructure cost-efficiency project. Figures were checked against billing data and reviewed by a separate agent that pushed back on the assumptions behind them — baseline choice and like-for-like comparability. The result is a set of savings figures with clearly labelled ranges and a short list of items still to verify, plus a method the client can repeat each month.

"Bringing the right context into each task through agents reduced hallucinations and raised output quality; AI peer reviews lightened the load on developers and lifted quality early, with faster turnaround."

What the client measured

The client’s own assessment of outcomes, reported at the close of the engagement:

Code and artefacts are now consistent across teams regardless of individual background or seniority. Problems are being caught before merge, where the cost of fixing them is lowest. Routine review work and environment support are increasingly handled by agents, freeing senior engineers for work that requires judgement. New developers reach productivity faster. Large, repetitive code changes now move at a pace the team considers sustainable, with consistent results. And because every output is grounded in connected systems with human sign-off, the results are reliable enough to build on.

The client is continuing to extend the capability, and has indicated that the patterns developed during the engagement are now forming the basis of its broader AI engineering strategy.

"I'm using it when a customer success manager asks if something can be done custom. In the past I'd reply with a simple answer; now I'm trying to give them a proof-of-concept."

*Client details have been anonymised.

Our latest posts

Got a High-Performance Team?

CLEVERBIT SCORE

Let’s find out.
Complete your Team’s Performance & AI Scorecard.
It only takes 2 minutes.

Scroll to Top
```