Case Study

AI-DLC at HBLab: Faster Delivery and 24/7 Cloud Operations with CloudThinker

AI-assisted coding made HBLab write code faster, and the bottleneck moved downstream to review, QC and 24/7 operations across many customer environments. HBLab applied CloudThinker at three points in its AI-DLC (review, operate, report) and saved thousands of hours while keeping a human approval gate on production.

casestudyvibecodingaicodereviewagenticopsmanagedcloudcostopshealthcheckcloudthinker
Cover Image for AI-DLC at HBLab: Faster Delivery and 24/7 Cloud Operations with CloudThinker

AI-DLC at HBLab: Faster Delivery and 24/7 Cloud Operations with CloudThinker

A developer at a software delivery company opens three pull requests before lunch, each written with an AI assistant in a fraction of the time it used to take. The reviewer has the same afternoon they always had. QC has the same test plan. The operations engineer on call tonight looks after a dozen customer environments, and any defect that slips through all of that ends up on their phone.

That's the shape of the problem HBLab faced. HBLab is a software delivery company whose teams ship continuously across many customer environments at once. AI-assisted development, the "VibeCoding" era, changed that work: code now gets written faster than ever. Writing was never the real constraint, though. As output climbed, the pressure moved downstream, to review, to QC, and to keeping production healthy without burning out the people who run it.

This is how HBLab made CloudThinker its standard AgenticOps platform: catching issues before they reach QC or production, running managed cloud operations 24/7 as a partnership, and automating the routine work behind every customer environment. The result freed thousands of hours and made the infrastructure HBLab operates more secure and more effective.


About HBLab

HBLab builds and delivers software (web, mobile and cloud) for a broad base of customers. Its engineers ship continuously across many client projects and environments, so quality and operational reliability are never one team's private concern. They're a promise renewed on every delivery, in every environment HBLab is responsible for.

That breadth is HBLab's strength. It's also what makes operating at speed hard: the faster the delivery engine runs, the more places a small mistake can land.


Why faster code didn't mean faster delivery

AI-assisted coding sped up how quickly HBLab's teams could produce working software. That speed-up is real, and it exposed something about delivery: the slow part was never typing the code. It was everything that has to be true after the code exists.

01Review

Falls behind

More changes, arriving faster, than human reviewers can read closely.

02Production

Ops absorbs it

Every slipped defect becomes a page for the operations team.

03Routine work

Gets skipped

Health, cost and performance checks are the first thing to go when people are stretched.

Where the VibeCoding speed-up landed. The bottleneck moved downstream of the code.

The traditional delivery model breaks under that kind of acceleration, for structural reasons:

Review capacity is fixed per person. More changes, arriving faster, meant more surface area to inspect. Human review didn't scale at the same rate, so defects that should have been caught early slipped toward QC, and sometimes all the way to production.

Downstream teams pay for upstream speed. Every customer environment HBLab ran had to stay healthy around the clock. The more the delivery engine accelerated, the more the operations team absorbed the cost: firefighting, context-switching, and a constant risk of burnout.

Routine operational work is the first thing dropped. Health checks, cost reports and performance reviews across many environments are high-volume, repetitive work. They get skipped when the team is stretched, and skipping them is how small problems grow into large ones.

Coverage thins out after hours. A lean team can't staff every environment at 3 a.m. with the same attention it gets at 3 p.m.

The goal wasn't to slow delivery down. It was to keep the speed the VibeCoding era gave HBLab, without paying for it in production incidents and exhausted engineers.


Where CloudThinker fits in the AI-DLC

HBLab didn't bolt automation onto the end of the pipeline. It mapped its AI-Driven Development Lifecycle (AI-DLC) end to end, from planning a change to operating it in production, and applied CloudThinker at the three points where the VibeCoding speed-up was doing the most damage: at review, before defects reach QC or production; at operate, keeping every environment healthy around the clock; and at report, turning routine health, cost and performance work into something that happens on its own.

01Plan and build

AI-assisted

Requirements, design and code, faster than ever.

02Review

AI Code Review

Bugs, security and performance issues caught before QC.

03Test and release

QC, then ship

QC validates far fewer preventable defects.

04Operate

24/7 managed

CloudThinker runs operations, humans approve production actions.

05Report

Automated

Health, cost and performance reports feed the next plan.

HBLab's AI-Driven Development Lifecycle. CloudThinker sits at review, operate and report; the report closes the loop back into planning.

The rest of this story walks those three touchpoints in order.

Review: catching issues before QC and production

The first change was to move quality left, to the moment code is written instead of the moment it breaks.

CloudThinker's AI Code Review runs on every pull request. It reads each change the way an experienced reviewer would, checking for correctness bugs, security issues and performance problems, and it does that on every change, consistently, without waiting for a human reviewer to free up. Issues that used to surface in QC or in production now get fixed while the code is still in review. The same review engine reached roughly 97% precision on FPT Cloud's production merge requests, evidence that the quality gate holds under production load. The code review benchmark shows how it handles the review bottleneck.

The effect compounds in a high-velocity delivery organization:

  • Delivery quality went up without slowing delivery down. The faster the team shipped, the more valuable a consistent, automatic quality gate became, because it scaled with the code instead of falling behind it.
  • The operations team stopped inheriting preventable defects. Fewer bugs reaching production meant fewer late-night incidents and less firefighting.

HBLab kept the velocity and got back the quality gate that velocity had outgrown.

Operate: one AgenticOps platform and a 24/7 partnership

Catching defects earlier solved half the problem. The other half was operating everything HBLab runs, reliably, at all hours.

HBLab standardized on CloudThinker as its AgenticOps platform, the single, consistent way its teams operate cloud infrastructure, and partnered with CloudThinker to run managed cloud operations 24/7. Instead of coverage that thinned out after business hours, HBLab's environments now get continuous operational attention. CloudThinker watches, triages and handles the routine work, while humans stay in control of anything that touches production.

01Environments

Customer clouds

  • AWSAWSCustomer accounts
  • AzureAzureCustomer subscriptions
  • Google CloudGoogle CloudCustomer projects

Many environments, one standard

02Operations

CloudThinkerCloudThinker

  • Watch and triageAround the clock
  • Health checksScheduled and on demand
  • CostOps reportsSpend and optimization
  • Performance reportsRecurring, not ad hoc

Same platform everywhere

03Judgment

HBLab engineers

  • Approve production actionsHuman gate on anything that touches prod
  • Act on findingsWith evidence attached
  • Report to customersHealth, cost and performance

Agents handle volume, people decide

The operating model HBLab standardized on across customer environments.

That last point matters. Production-affecting actions keep a human approval gate. CloudThinker handles the volume; HBLab's engineers keep the judgment. The result is a 24/7 operating model a lean team could never have staffed on its own, and one standard that works the same way across every customer environment.

Report: health, cost and performance without the afternoon of work

With CloudThinker as the standard platform, HBLab automated the operational work that used to be skipped for lack of time. For every customer environment, CloudThinker now runs it on a schedule and on demand:

  • System health checks, continuous and across the stack, so problems are seen early instead of discovered during an incident.
  • CostOps reports, turning cloud spend from a quarterly surprise into a managed, reviewed number.
  • Performance reports, the recurring reviews that used to take a specialist's afternoon, now produced automatically.

Automating this work saved thousands of hours of repetitive effort. More importantly, it let HBLab deliver reporting and operational rigor it had never been able to sustain manually across so many environments. The infrastructure HBLab operates for its customers became more secure and more effective, because the work that keeps it that way now happens consistently, everywhere.


What a morning health check finds

The walkthrough below is illustrative. Resource names and values are made up, and it isn't a record of a specific HBLab customer. It shows the kind of finding a scheduled health check turns up before it becomes an incident.

The signal. The overnight check on one customer environment notes that free storage on a production PostgreSQL instance on RDS has dropped steadily for four days. Nothing has alarmed yet; at this rate it will in about a week.

The readings. Steady storage loss on Postgres has a few common causes, and each points somewhere different:

  • Real data growth. Table sizes would grow in step with the storage curve.
  • Log retention. Log volume would rise, typically after a parameter change.
  • Retained WAL. An inactive replication slot (often left behind by a decommissioned CDC or replication job) stops Postgres from recycling WAL, so transaction logs pile up.

The checks. First, the storage trend next to transaction log usage:

aws cloudwatch get-metric-statistics \
  --namespace AWS/RDS \
  --metric-name TransactionLogsDiskUsage \
  --dimensions Name=DBInstanceIdentifier,Value=orders-prod \
  --start-time 2026-09-26T00:00:00Z --end-time 2026-09-30T00:00:00Z \
  --period 3600 --statistics Maximum

Then the replication slots on the instance itself:

SELECT slot_name,
       active,
       pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal
FROM pg_replication_slots;

What came back (illustrative). Table sizes are flat, so it isn't data growth. Transaction log usage climbs in a straight line that matches the storage loss. One slot is inactive and holding tens of gigabytes of WAL. It belongs to a CDC connector that was switched off the previous week.

The finding and the gate. The report names the slot, the retained size, the connector it belonged to and the date it stopped. The proposed fix is to drop the slot once the owner confirms the connector is gone for good. Dropping a slot is a production write, so it waits for an HBLab engineer to approve it. Nobody gets paged, and the storage alarm never fires.


The outcome: a new operating model

Operational dimension Before After
Delivery quality gate Manual review, defects reach QC/prod AI review on every change, caught pre-QC
Operations coverage Business hours, reactive 24/7 managed cloud, human-approved actions
Routine reporting Manual or skipped Automated health, cost and performance
Operations team load Firefighting, burnout risk Focused on high-value work
Infrastructure posture Uneven across environments More secure and effective, consistently

The deeper difference is where speed and stability meet. In the old model they traded against each other: every gain in delivery speed was paid for by review depth and operations load. With review, operations and reporting on one platform, the delivery engine kept the speed the VibeCoding era gave it, and the work that keeps production healthy scaled with it instead of lagging behind.


Gated by design

Running operations for customers means every action has to be explainable to the customer. CloudThinker acts within Auto Mode, so the team decides what runs on its own, such as sending the weekly cost report, and what waits for an engineer's approval, such as dropping a replication slot. Every action carries the evidence and the reasoning.

Getting started

Connect one customer environment read-only, then ask from chat:

"Run a health check across this environment and list anything trending toward an alarm in the next two weeks, with the evidence for each. Propose only, change nothing."

"Build last month's cost report for this account, broken down by service, and flag the three biggest changes from the month before."

"Review every pull request on this repository for correctness, security and performance issues before it goes to QC."


Related reading

Conclusion

The VibeCoding era made writing software faster. On its own, it didn't make delivering and operating that software safer; the bottleneck moved downstream. By making CloudThinker its standard AgenticOps platform, HBLab closed that gap: quality moved left to code review, operations moved to a 24/7 partnership, and the routine work that keeps infrastructure healthy became automatic. The delivery team ships, the operations team isn't burning out, and the infrastructure behind every customer gets more secure and more effective over time.

To see how CloudThinker does this, explore AI Code Review and 24/7 Managed Cloud, or talk to our team.