Patching, certificate and secret rotation, scaling and storage done on schedule.
Patching, certificate and secret rotation, storage expansion and scaling are the work that keeps production up and never makes the roadmap. Agents run it on the schedule and policy you set, check each change before and after, and hand your team a record of what ran and why.
Monthly patching: 212 instances, 3 accounts
Compliance report sent to #ops-maintenance
[The work behind every maintenance window]
01The manual work
Maintain
Engineers track expiry dates in spreadsheets, patch in late-night windows and grow disks after the alert, not before.
02The agent handoff
Done on time
Frontier agents plan each task, check the system before and after, run it in safe waves and roll back if health checks fail.
03Your engineers’ role
Approve
Set the schedule, the blast radius and what needs a person. Review the exceptions, not the routine.
[Where CloudThinker fits]
01Fleet
Wherever your hosts run
No change to your tooling
02Runbooks and change
Your system of record
Change process stays put
03Routine work
CloudThinker
Runs inside your change window
04Response
Done, then reported
Each wave runs on approval
[Example scenario]
A 12-person ops team running 212 EC2 instances across 3 AWS accounts, plus 40 TLS certificates and 90 service secrets. A critical OpenSSL patch landed on Thursday.
Pre-checks runAgent
Confirms the patch baseline, takes EBS snapshots and skips 2 hosts running finance batch jobs until they finish.
Wave 1 of 4Agent
Patches 25% of instances through Systems Manager, reboots, then waits for load balancer health checks to pass.
One host failsSignal
ip-10-2-14-8 does not come back after reboot. Health check fails for 5 minutes.
Rolled back, rollout continuesAgent
Restores the host from its snapshot, pulls it from the wave, opens OPS-774 with the console log, and continues with the healthy hosts.
Rotation and capacityAgent
Renews 3 certificates due in 14 days, rotates 6 database secrets with zero failed logins, and grows a log volume at 81% full.
Team reviews one exceptionYour team
The on-call engineer reads the morning summary, checks OPS-774 and approves a rebuild of the failed host.
CloudThinker00:45
Starting monthly patching: 212 instances in 4 waves. Snapshots done. Skipping 2 hosts with active batch jobs until 02:30.
CloudThinker01:40
ip-10-2-14-8 failed to boot after patching. Restored from snapshot and removed from this run. Opened OPS-774 with the console log. Other waves continuing.
CloudThinker03:40
Done. 209 patched and healthy, 3 rescheduled. Renewed 3 certificates, rotated 6 secrets, expanded logs-vol-02 to 500 GB. Compliance report attached.
On-call engineer08:30
Read the log on OPS-774. Approved the rebuild.
An illustrative example. Team, systems and times are representative, not a specific customer.
[Frontier resolution agents]
Agents run day-2 work on the same policy every time, so maintenance stops depending on who remembers what, and every change leaves a record your auditors can read.
[What changes]
| Moment | Today | With frontier agents |
|---|---|---|
| Patch night | Engineers awake until 4am | Runs in waves, reviewed over coffee |
| Certificate expiry | Found when the site breaks | Renewed weeks ahead, checked after |
| Secret rotation | Skipped because it might break things | Rotated on schedule with each service verified |
| Disk space | Expanded after the 95% alert | Grown early from the usage trend |
| Audit evidence | Screenshots collected by hand | A report from every run |
[Integrations]
[Adoption path]
The rollout follows the four phases of the AWS Cloud Adoption Framework, so it fits the plan your cloud team already runs.
01Envision
Start with one runbook
Pick the task your team dreads most, often patching or certificate renewal, and let agents plan it in dry-run mode.
02Align
Set windows and limits
Agree maintenance windows, wave sizes, health checks and which failures stop the run.
03Launch
Add the other routines
Secret rotation, storage, scaling and log rotation join on the same schedule and the same report.
04Scale
Make it the default
New accounts and services inherit the routines on day one. Your team only sees the exceptions.
[Customer proof]
A Vietnamese consumer-finance company with 800+ branches encoded its runbooks once and applied them across accounts and environments.
[AWS guidance]
[Trust and control]
[Questions]
[Go deeper]
Start with one runbook in dry-run mode. See the plan agents produce before you let them run a single change.

Up to $200K in AWS credits
Applied to your own AWS account.

AWS AI Services Competency
Validated for Agentic AI Consulting.

Covered 24/7, on your approval
Under HIPAA, GDPR and SOC 2 controls.