Question 12
For a machine learning API serving models with inconsistent response times, which monitoring strategy provides the most actionable insights?
Track only HTTP response status codes
Model inference time + request queue depth + resource utilization + error rate tracking
Monitor only successful prediction accuracy
Focus exclusively on network bandwidth usage