
The last working Friday before Christmas has its own smell: mince pies in the kitchen, half the team already on leave, and a backlog nobody wants to touch. It is also the week the internet gets busy. Retail traffic climbs, sign-up forms get hammered, and the person who last ran the failover script is somewhere with one bar of signal.
Most UK engineering teams get through the holidays without incident. Those that don't tend to fail for dull reasons: an alert that fired into a channel nobody was watching, an autoscaling group that hit its ceiling on Boxing Day, a certificate that expired while everyone was eating leftovers. Five checks will catch most of it, and none of them takes more than an afternoon.
Check one: alerts that reach a person
Start with the paging path, not the metrics. An alert is only useful if it lands on a device belonging to someone who can act, and holiday rotas quietly break that assumption. Overrides get set in the on-call tool and forgotten; a team of eight becomes a rota of two. The person who knows how the payments service behaves under load is skiing, and the person covering has never seen a queue depth graph before.
Send a test alert end to end. Trigger it through the monitoring pipeline rather than the test button in the dashboard, and confirm it rings a phone instead of filling an inbox nobody checks. Then look at what fired over the last month and ask which of those alerts anyone actually acted on. Mute the rest deliberately, with a maintenance window and a note, so the on-call engineer doesn't learn to swipe notifications away.
A three-minute alert audit
Open your alerting tool and sort by frequency. For each of the top ten alerts, write down the action it should prompt. If the action is "look at it in the morning", it is not a page; it is a ticket. If the action is "restart the worker", make sure the runbook says which worker and how to check that the restart worked. Alerts without actions are noise, and noise is how real pages get missed.
Two smaller things while you are there. Confirm your log retention covers the whole break: if something surfaces in mid-January, you want December's logs still searchable. And make sure every alert includes a runbook link, even if the runbook is three lines of shell.
Check two: where you will run out of headroom
Autoscaling has ceilings, and those ceilings are usually set by quotas you requested long ago, when the numbers were smaller. A holiday spike doesn't just need CPU; it needs IP addresses, database connections, queue throughput and API rate limits that must exist before the load arrives. The worst time to discover a quota limit is when the scaling event is already happening and the support ticket queue is staffed by a skeleton crew.
- Compute: instance and node limits per region, the cluster autoscaler's maximum node count, and whether spot capacity is likely to be available on a public holiday. Spot markets can dry up when everyone else is also scaling.
- Networking: free addresses in each subnet, NAT gateway throughput, load balancer limits, DNS query quotas. A subnet with twelve free IPs cannot absorb a hundred new pods.
- Data: database connection limits, storage headroom, replica lag under write-heavy load, and what happens when a queue fills. Decide now whether a full queue should drop messages, block writes, or page someone.
- Managed services: concurrency limits on serverless functions, provisioned throughput, and rate limits on any third-party API you call on the critical path. If your checkout calls a tax service, know its quota and its holiday support hours.
- Account level: spending limits and region restrictions that could block a scale-up, plus any quota increase that needs a support ticket to clear. Some cloud providers require business justification for large increases, so include your expected peak in the request.
Quota requests take days, sometimes longer if the reviewer is also on leave. Submit them now, with a buffer you would normally consider excessive. It is cheaper than an outage. If a quota increase is refused, decide the graceful degradation path: which non-critical feature you will turn off first to protect checkout, login, or whatever pays the bills.
Check three: freeze changes, not your ability to fix things
A change freeze is a sensible default for the last fortnight of December, but an absolute freeze without an exception path just pushes deployments underground. Agree the rules in writing: what counts as an emergency, who can approve it, and how quickly that approval can be obtained at 7pm on Christmas Eve. A freeze that nobody can override is a freeze that gets ignored when a real problem appears.
A freeze is a promise that the system will stay still. It is not a promise that the system will stay healthy.
Then make the rollback path boring. Know which release is currently in production, keep the previous image or build artefact tagged and available, and verify that the feature flag you would switch off still works. Finish any half-applied database migration before the freeze rather than leaving a schema only part of the code understands. Experimental work can sit on a branch until January; it will not improve over the break.
Write the exception process where people will find it: pinned in the on-call channel, linked from the runbook, and included in the handover notes. Include the name and phone number of the person who can approve a hotfix, not just a role title. If that person is unavailable, name a backup.
Check four: an on-call rota that survives annual leave
Write the rota down with names, not team names. "Platform team" is not a person who can be woken at 3am. Every shift should have a primary and a secondary, and both should know they are on. Check that the rota tool's overrides match the spreadsheet, because the two drift apart the moment someone swaps a shift in a group chat.
- Primary and secondary: two names per shift, with a clear escalation path if the primary does not answer within ten minutes.
- Time zones and travel: note who is abroad, who has unreliable signal, and who is driving on a particular day. A rota that assumes instant response from someone on a ferry is not a rota.
- Access: confirm every on-call engineer can reach the cloud console, the CI/CD system, the database, and the incident channel. Check that their MFA device works and their break-glass credentials have not expired.
- Handover: require a short written handover at each shift change: what is broken but known, what is being watched, and what should not be touched.
Finally, make the first hour of an incident easy. A one-page guide with links to dashboards, the status page, and the customer comms template saves more time than any amount of heroic debugging. If the incident is severe, the on-call engineer should be able to declare it, pull in help, and communicate without hunting for permissions.
Check five: the quiet expiries
Certificates, domains, API tokens and backup schedules do not care about your change freeze. They expire on their own timetable, and often at the least convenient moment. A TLS certificate that lapses on Christmas Day will take down every customer-facing endpoint, and the fix will be complicated by the fact that the person who owns the DNS record is unreachable.
- TLS certificates: list every certificate, its expiry date, and the renewal owner. Automate renewal where possible, and test the renewal path rather than assuming it works.
- Domains and DNS: check registration expiry and auto-renew settings. Confirm that DNS changes can be made by more than one person.
- Secrets and tokens: rotate anything due before mid-January now, and check that service accounts do not have credentials that expire over the break.
- Backups: verify that backups are running, encrypted, and restorable. A backup that has never been restored is a hope, not a plan.
Do one restore test before the freeze. Restore a small database into a staging environment and time it. If the restore takes six hours, you need to know that before an incident, not during one.
Extra: run a 90-minute pre-freeze game day
Book a single session with the people who will be on call and walk through three failure scenarios. Keep it short and concrete:
- What happens if the primary database fails over during peak traffic? Who gets paged, what do they check first, and how do they confirm the failover completed?
- What happens if a third-party API starts returning 429s? Which feature degrades, and how do you communicate that to customers?
- What happens if the person on call cannot access the cloud console? Where is the break-glass process, and has it been tested this quarter?
The goal is not to solve everything. It is to expose the gaps while there is still time to fix them. If the answer to any scenario is "I think someone has a script", that is your next task.
Extra: the holiday incident one-pager
Keep a single page at the top of the on-call channel. It should contain the escalation path, the status page URL, the customer comms template, and the three largest risks you identified above. Add a short line for each risk: what it looks like, what to check first, and who owns the fix. That page will be read at the worst possible moment, so make it scannable and free of jargon.
None of these checks is glamorous. They are the equivalent of checking the boiler before a cold snap: filters, pressure, thermostat. Do them now, while the office is still half full and the support queue is still quiet, and the holiday freeze becomes what it should be: a quiet fortnight where nothing pages anyone until January.
Photo: Altaf Shah / Pexels


