A female engineer using a laptop while monitoring data servers in a modern server room.

Photo by Christina Morillo on Pexels

An EC2 instance can remain available while the application running on it fails to serve customers. AWS’s application-level health checks narrow that blind spot by monitoring an HTTP or HTTPS endpoint and allowing Auto Scaling to replace instances whose applications become unhealthy.

At 2:07 a.m., the console can still show green checks while customer requests time out. Those checks may accurately report that the instance is running and reachable from AWS’s infrastructure perspective. They do not necessarily prove that the web server is responding, the application has finished starting, or a critical dependency is working.

What the green check actually proves

Infrastructure health checks answer infrastructure questions. They can identify failures involving the underlying host, networking, power, or an instance that has stopped responding at the expected level.

That evidence matters. It simply covers a narrower failure boundary than customers experience.

A running virtual machine can contain a dead process. An application can accept TCP connections but stall every request. A deployment can leave the service trapped in a restart loop. Memory exhaustion can make responses unusably slow without immediately making the instance disappear. In each case, the infrastructure may remain healthy while the customer-facing path has failed.

The distinction is easy to lose during an incident because dashboards compress several layers into one reassuring color. “Healthy instance” sounds broader than the underlying test warrants. Operators need to read it as a precise statement: the infrastructure check passed.

Application reachability requires separate evidence.

AWS moves health checks closer to the customer path

AWS has added application-level health checks to EC2. Customers can monitor HTTP or HTTPS endpoints, giving the platform a signal from the application layer rather than relying only on instance availability.

The practical change is in how Auto Scaling can respond. When an application endpoint becomes unhealthy, an Auto Scaling group can replace the affected instance. That creates a recovery path for failures where the machine continues running but the service on it does not.

The value depends on the endpoint being tested. A shallow endpoint that returns `200 OK` whenever the process is alive may miss the same failures operators want to catch. A deeper check can test whether the application is ready to handle meaningful work, but every added dependency introduces another decision: should failure trigger instance replacement?

That choice needs care. If every instance depends on the same unavailable database, replacing the fleet will not restore service. It may add startup load while the shared dependency is already struggling. Health checks should distinguish an instance-specific fault from a system-wide outage whenever the architecture allows it.

The check also needs enough time for normal startup. Marking a new instance unhealthy before migrations, caches, or application initialization finish can create a replacement loop. Thresholds, grace periods, and endpoint behavior matter as much as enabling the feature.

Availability needs evidence from more than one layer

The strongest monitoring design uses several independent views of the service.

Instance checks show whether the computing foundation is available. Application checks show whether a particular endpoint responds. Load balancer health checks show whether a target can receive traffic through that path. External probes show what a request looks like from outside the immediate AWS environment.

None of those signals alone proves that a customer can complete a purchase, upload a file, or retrieve an account record. An HTTP health endpoint may return successfully while authentication, payments, or another critical workflow fails.

That is why teams should map each green indicator to the claim it supports. If a dashboard tile says “application healthy,” the underlying test should exercise enough of the application to justify the label. If it only confirms that a process returns a response, label it accordingly.

This is the same operational trap examined in Friday Green, Friday Broken: status indicators are useful only when their scope is understood. A monitoring panel can be entirely accurate and still leave the most important question unanswered.

What operators should change before the next timeout

Start by documenting the current health-check chain from EC2 through the load balancer and application. Record what each check sends, what response counts as healthy, how many failures trigger action, and which system takes that action.

Then test failure modes deliberately in a safe environment. Stop the application process while leaving the instance running. Make the endpoint hang. Return an error. Block a dependency used by the readiness check. Confirm that monitoring changes state, traffic stops reaching the bad target, and Auto Scaling behaves as intended.

Watch for replacement storms during those tests. A health check that correctly detects a shared dependency failure can still trigger the wrong recovery action. Alerting may be appropriate where replacement would only recycle healthy machines into the same outage.

Finally, add a customer-path probe for the operation that matters most. Keep it deterministic and avoid destructive transactions. The goal is to establish whether the service can perform useful work, from a location and route that resemble real access.

The next time the console remains green at 2:07 a.m., the team should know exactly what that green check proves, what it leaves untested, and which signal can confirm whether customers are getting through.

Comments

No comments yet.