QuestionQ58

Monitoring and Logging

A mobile application operating on eight Amazon EC2 instances depends on a third-party API endpoint. The third-party service has a high failure rate because its capacity is limited, which is expected to be fixed in a few weeks.

Meanwhile, the mobile application developers have implemented a retry mechanism and are logging failed API requests. A DevOps engineer must automate application-log monitoring and count the specific error messages. If there are more than 10 errors in a 1-minute window, the system must send an alert.

How can these requirements be met with MINIMAL management overhead?

Explanation

Amazon CloudWatch Logs metric filters can match specific error messages as application logs are ingested and transform each match into a CloudWatch metric value. CloudWatch aggregates the metric every minute, and a CloudWatch alarm can alert when the one-minute error count is greater than 10. Using the CloudWatch agent to send the instance application logs to CloudWatch Logs avoids maintaining custom polling scripts or EventBridge-based counting logic.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!