A mobile application operating on eight Amazon EC2 instances depends on a third-party API endpoint. The third-party service has a high failure rate because its capacity is limited, which is expected to be fixed in a few weeks.
Meanwhile, the mobile application developers have implemented a retry mechanism and are logging failed API requests. A DevOps engineer must automate application-log monitoring and count the specific error messages. If there are more than 10 errors in a 1-minute window, the system must send an alert.
How can these requirements be met with MINIMAL management overhead?
Community Discussion
No comments yet. Be the first to start the discussion!
Community Discussion