When GitHub starts failing, spend the first five minutes confirming scope, preserving evidence and controlling retries. Pause merges, releases and automation until you know whether the failure sits with GitHub, your network, your credentials or a single repository.
At 09:00, the first symptom may look harmless: a pull request page times out, `git push` stalls, or an Actions job never starts. Seconds later, someone retries the deployment while another person rotates a token. If the incident affects several GitHub services, those well-meant actions can erase useful evidence, create duplicate jobs and turn one external failure into several internal mysteries.
00:00 to 01:00: Confirm the symptom from two paths
Record the time in UTC and copy the exact error before refreshing anything. Capture the command, endpoint or page involved, along with the repository, workflow or organization affected. A screenshot helps, but plain text is easier to search later.
Run one low-impact check through the failing path, then compare it with a second path. If the website fails, test a read-only CLI or API request. If `git push` fails, try a fetch or a request against another repository you can already access. Ask one colleague on a different network to check the same service.
The goal is classification, not diagnosis. You want to know whether the evidence points toward:
- one user, repository or organization;
- your office network, VPN, DNS or identity provider;
- one GitHub feature, such as Actions;
- several GitHub surfaces failing together.
Avoid destructive tests. Do not rotate credentials, rerun every workflow, rebase branches or force-push while the failure remains unexplained. Those actions change the system you are trying to observe.
01:00 to 02:00: Preserve evidence before retries overwrite it
Create a short incident note in a system that does not depend on GitHub. Record the first observed failure, exact error text, affected services, successful comparison checks and the names of any jobs already in progress.
Save request IDs and response headers when available. For CLI failures, keep the command output and exit code. For Actions, note the workflow name, run identifier, commit SHA and its last visible state. For API calls, record the endpoint, method, status code and timestamp, while keeping tokens and credentials out of the note.
Be careful with screenshots. They can reveal private repository names, customer data, internal branch names or browser extensions. Crop them before sharing outside the incident channel.
Evidence matters because a broad outage may recover before your team understands what happened. GitHub said its August 17 outage lasted 7 hours and 47 minutes and affected authentication, Actions, APIs, pull requests, issues, Copilot and the main website. When several dependent services fail together, a single error message rarely describes the full boundary.
02:00 to 03:30: Stop automation from multiplying the damage
Freeze actions that can create duplicates or conflicting state. Pause manual deployments. Tell engineers to stop rerunning failed workflows until an incident lead approves retries. If release automation calls GitHub APIs, package registries or Actions, treat an ambiguous timeout as an unknown outcome rather than a clean failure.
That distinction matters. A request can complete on the remote side even when the client never receives confirmation. Repeating it may create a second release, duplicate comment, repeated infrastructure change or competing deployment.
Do not cancel every running job by default. Cancellation is another state-changing operation, and the request itself may fail midway. First identify which jobs are safe to leave alone, which could spend money indefinitely, and which could alter production if GitHub resumes.
Assign one person to make retry decisions. Everyone else reports observations without experimenting independently. This follows the same containment principle described in The Automation Broke Before Stand-Up: uncontrolled recovery attempts often produce a second incident that is harder to reconstruct than the first.
03:30 to 05:00: Set ownership and the next checkpoint
Name an incident lead, a technical investigator and a communicator. In a small team, one person may hold two roles, but retry authority should remain clear.
Post a compact status update:
- GitHub symptoms began at the recorded UTC time.
- The confirmed affected paths are listed, with unknowns marked separately.
- Merges, deployments or workflow reruns are paused.
- Production impact is confirmed, suspected or absent.
- The next update will arrive at a specific time.
Keep facts separate from analysis. “API requests returned a specific status code” is an observation. “GitHub authentication is broken” is a hypothesis unless broader checks support it. “No production impact observed” also has limits; say what was checked.
Set the next checkpoint five or ten minutes ahead. Until then, run only targeted, read-only checks that can distinguish competing explanations. Keep one untouched failing example for later comparison. If service returns, test recovery with a low-risk read before restarting queued writes, merges or deployments.
The fifth minute should end with fewer people clicking buttons, a preserved record of the first failure and one named person deciding what happens next.
Comments
No comments yet.