AI-DLC at HBLab: Faster Delivery and 24/7 Cloud Operations with CloudThinker
A developer at a software delivery company opens three pull requests before lunch, each written with an AI assistant in a fraction of the time it used to take. The reviewer has the same afternoon they always had. QC has the same test plan. The operations engineer on call tonight looks after a dozen customer environments, and any defect that slips through all of that ends up on their phone.
That's the shape of the problem HBLab faced. HBLab is a software delivery company whose teams ship continuously across many customer environments at once. AI-assisted development, the "VibeCoding" era, changed that work: code now gets written faster than ever. Writing was never the real constraint, though. As output climbed, the pressure moved downstream, to review, to QC, and to keeping production healthy without burning out the people who run it.
This is how HBLab made CloudThinker its standard AgenticOps platform: catching issues before they reach QC or production, running managed cloud operations 24/7 as a partnership, and automating the routine work behind every customer environment. The result freed thousands of hours and made the infrastructure HBLab operates more secure and more effective.
About HBLab
HBLab builds and delivers software (web, mobile and cloud) for a broad base of customers. Its engineers ship continuously across many client projects and environments, so quality and operational reliability are never one team's private concern. They're a promise renewed on every delivery, in every environment HBLab is responsible for.
That breadth is HBLab's strength. It's also what makes operating at speed hard: the faster the delivery engine runs, the more places a small mistake can land.
Why faster code didn't mean faster delivery
AI-assisted coding sped up how quickly HBLab's teams could produce working software. That speed-up is real, and it exposed something about delivery: the slow part was never typing the code. It was everything that has to be true after the code exists.
01Review
Falls behind
More changes, arriving faster, than human reviewers can read closely.
02Production
Ops absorbs it
Every slipped defect becomes a page for the operations team.
03Routine work
Gets skipped
Health, cost and performance checks are the first thing to go when people are stretched.
The traditional delivery model breaks under that kind of acceleration, for structural reasons:
Review capacity is fixed per person. More changes, arriving faster, meant more surface area to inspect. Human review didn't scale at the same rate, so defects that should have been caught early slipped toward QC, and sometimes all the way to production.
Downstream teams pay for upstream speed. Every customer environment HBLab ran had to stay healthy around the clock. The more the delivery engine accelerated, the more the operations team absorbed the cost: firefighting, context-switching, and a constant risk of burnout.
Routine operational work is the first thing dropped. Health checks, cost reports and performance reviews across many environments are high-volume, repetitive work. They get skipped when the team is stretched, and skipping them is how small problems grow into large ones.
Coverage thins out after hours. A lean team can't staff every environment at 3 a.m. with the same attention it gets at 3 p.m.
The goal wasn't to slow delivery down. It was to keep the speed the VibeCoding era gave HBLab, without paying for it in production incidents and exhausted engineers.
Where CloudThinker fits in the AI-DLC
HBLab didn't bolt automation onto the end of the pipeline. It mapped its AI-Driven Development Lifecycle (AI-DLC) end to end, from planning a change to operating it in production, and applied CloudThinker at the three points where the VibeCoding speed-up was doing the most damage: at review, before defects reach QC or production; at operate, keeping every environment healthy around the clock; and at report, turning routine health, cost and performance work into something that happens on its own.
01Plan and build
AI-assisted
Requirements, design and code, faster than ever.
02Review
AI Code Review
Bugs, security and performance issues caught before QC.
03Test and release
QC, then ship
QC validates far fewer preventable defects.
04Operate
24/7 managed
CloudThinker runs operations, humans approve production actions.
05Report
Automated
Health, cost and performance reports feed the next plan.
The rest of this story walks those three touchpoints in order.
Review: catching issues before QC and production
The first change was to move quality left, to the moment code is written instead of the moment it breaks.
CloudThinker's AI Code Review runs on every pull request. It reads each change the way an experienced reviewer would, checking for correctness bugs, security issues and performance problems, and it does that on every change, consistently, without waiting for a human reviewer to free up. Issues that used to surface in QC or in production now get fixed while the code is still in review. The same review engine reached roughly 97% precision on FPT Cloud's production merge requests, evidence that the quality gate holds under production load. The code review benchmark shows how it handles the review bottleneck.
The effect compounds in a high-velocity delivery organization:
- Delivery quality went up without slowing delivery down. The faster the team shipped, the more valuable a consistent, automatic quality gate became, because it scaled with the code instead of falling behind it.
- The operations team stopped inheriting preventable defects. Fewer bugs reaching production meant fewer late-night incidents and less firefighting.
HBLab kept the velocity and got back the quality gate that velocity had outgrown.
Operate: one AgenticOps platform and a 24/7 partnership
Catching defects earlier solved half the problem. The other half was operating everything HBLab runs, reliably, at all hours.
HBLab standardized on CloudThinker as its AgenticOps platform, the single, consistent way its teams operate cloud infrastructure, and partnered with CloudThinker to run managed cloud operations 24/7. Instead of coverage that thinned out after business hours, HBLab's environments now get continuous operational attention. CloudThinker watches, triages and handles the routine work, while humans stay in control of anything that touches production.
01Environments
Customer clouds
AWSCustomer accounts
AzureCustomer subscriptions
Google CloudCustomer projects
Many environments, one standard
02Operations
CloudThinker
- Watch and triageAround the clock
- Health checksScheduled and on demand
- CostOps reportsSpend and optimization
- Performance reportsRecurring, not ad hoc
Same platform everywhere
03Judgment
HBLab engineers
- Approve production actionsHuman gate on anything that touches prod
- Act on findingsWith evidence attached
- Report to customersHealth, cost and performance
Agents handle volume, people decide
That last point matters. Production-affecting actions keep a human approval gate. CloudThinker handles the volume; HBLab's engineers keep the judgment. The result is a 24/7 operating model a lean team could never have staffed on its own, and one standard that works the same way across every customer environment.
Report: health, cost and performance without the afternoon of work
With CloudThinker as the standard platform, HBLab automated the operational work that used to be skipped for lack of time. For every customer environment, CloudThinker now runs it on a schedule and on demand:
- System health checks, continuous and across the stack, so problems are seen early instead of discovered during an incident.
- CostOps reports, turning cloud spend from a quarterly surprise into a managed, reviewed number.
- Performance reports, the recurring reviews that used to take a specialist's afternoon, now produced automatically.
Automating this work saved thousands of hours of repetitive effort. More importantly, it let HBLab deliver reporting and operational rigor it had never been able to sustain manually across so many environments. The infrastructure HBLab operates for its customers became more secure and more effective, because the work that keeps it that way now happens consistently, everywhere.
What a morning health check finds
The walkthrough below is illustrative. Resource names and values are made up, and it isn't a record of a specific HBLab customer. It shows the kind of finding a scheduled health check turns up before it becomes an incident.
The signal. The overnight check on one customer environment notes that free storage on a production PostgreSQL instance on RDS has dropped steadily for four days. Nothing has alarmed yet; at this rate it will in about a week.
The readings. Steady storage loss on Postgres has a few common causes, and each points somewhere different:
- Real data growth. Table sizes would grow in step with the storage curve.
- Log retention. Log volume would rise, typically after a parameter change.
- Retained WAL. An inactive replication slot (often left behind by a decommissioned CDC or replication job) stops Postgres from recycling WAL, so transaction logs pile up.
The checks. First, the storage trend next to transaction log usage:
aws cloudwatch get-metric-statistics \
--namespace AWS/RDS \
--metric-name TransactionLogsDiskUsage \
--dimensions Name=DBInstanceIdentifier,Value=orders-prod \
--start-time 2026-09-26T00:00:00Z --end-time 2026-09-30T00:00:00Z \
--period 3600 --statistics Maximum
Then the replication slots on the instance itself:
SELECT slot_name,
active,
pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal
FROM pg_replication_slots;
What came back (illustrative). Table sizes are flat, so it isn't data growth. Transaction log usage climbs in a straight line that matches the storage loss. One slot is inactive and holding tens of gigabytes of WAL. It belongs to a CDC connector that was switched off the previous week.
The finding and the gate. The report names the slot, the retained size, the connector it belonged to and the date it stopped. The proposed fix is to drop the slot once the owner confirms the connector is gone for good. Dropping a slot is a production write, so it waits for an HBLab engineer to approve it. Nobody gets paged, and the storage alarm never fires.
The outcome: a new operating model
| Operational dimension | Before | After |
|---|---|---|
| Delivery quality gate | Manual review, defects reach QC/prod | AI review on every change, caught pre-QC |
| Operations coverage | Business hours, reactive | 24/7 managed cloud, human-approved actions |
| Routine reporting | Manual or skipped | Automated health, cost and performance |
| Operations team load | Firefighting, burnout risk | Focused on high-value work |
| Infrastructure posture | Uneven across environments | More secure and effective, consistently |
The deeper difference is where speed and stability meet. In the old model they traded against each other: every gain in delivery speed was paid for by review depth and operations load. With review, operations and reporting on one platform, the delivery engine kept the speed the VibeCoding era gave it, and the work that keeps production healthy scaled with it instead of lagging behind.
Gated by design
Running operations for customers means every action has to be explainable to the customer. CloudThinker acts within Auto Mode, so the team decides what runs on its own, such as sending the weekly cost report, and what waits for an engineer's approval, such as dropping a replication slot. Every action carries the evidence and the reasoning.
Getting started
Connect one customer environment read-only, then ask from chat:
"Run a health check across this environment and list anything trending toward an alarm in the next two weeks, with the evidence for each. Propose only, change nothing."
"Build last month's cost report for this account, broken down by service, and flag the three biggest changes from the month before."
"Review every pull request on this repository for correctness, security and performance issues before it goes to QC."
Related reading
- Inside the FPT Cloud Evaluation: How CloudThinker AI Code Review Held the First Quality Gate: the code-review results behind the review stage, measured on production merge requests under enterprise security and compliance.
- What Takes Your Team Days, AI Does in Minutes: how AI review changed the review bottleneck.
- The Death of the Traditional SDLC: Why the VibeOps Era Needs a Guardrail: the industry shift behind the pressures in this story.
- 2:47 A.M. in Someone Else's Production: Inside CloudThinker's On-call AgenticOps Team: the detect, resolve and validate loop behind 24/7 coverage.
- How an Australian Telematics Provider Automates Day-to-Day Operations with CloudThinker: the same code-review, CostOps and health-check pattern in another operator.
Conclusion
The VibeCoding era made writing software faster. On its own, it didn't make delivering and operating that software safer; the bottleneck moved downstream. By making CloudThinker its standard AgenticOps platform, HBLab closed that gap: quality moved left to code review, operations moved to a 24/7 partnership, and the routine work that keeps infrastructure healthy became automatic. The delivery team ships, the operations team isn't burning out, and the infrastructure behind every customer gets more secure and more effective over time.
To see how CloudThinker does this, explore AI Code Review and 24/7 Managed Cloud, or talk to our team.
