QuestionQ18

AI Infrastructure Operations and Troubleshooting

An engineer needs to optimize the performance of an AI inference workload running on a Cisco UCS C-Series rack server. The workload has intermittent latency spikes, and Intersight system logs show frequent thermal-event entries. Which troubleshooting action must be performed first to resolve the performance issues on the Cisco UCS server?

Explanation

Thermal conditions can reduce server efficiency and lead to performance degradation, including latency spikes caused by thermal throttling. Thermal-event troubleshooting should begin with the physical cooling environment: verify ambient temperature and unobstructed airflow, and correct failed fan hardware. Cisco’s UCS fault guidance specifically directs administrators to check airflow and cooling and replace faulty fan modules for thermal faults.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!