Q
A company has a customer service agent that calls an Azure OpenAI model and sends traces to Azure Monitor Application Insights.
After a recent release, average response latency and monitoring costs have increased. Telemetry shows higher token usage per request and a large increase in ingested trace data.
You need to reduce monitoring costs, while retaining enough telemetry to analyze token usage and latency trends.
What should you do?
Question Info
Choose the Best Option
Click any option to instantly check if you're correct.
Explanation
Objective:
3.1 Analyze, monitor, and tune AI-powered business solutions
What This Item Tests:
Interpret telemetry data for performance and model tuning
Additional Reading:
Understand ongoing non-infrastructure costs of AI agents
Explore how to monitor with Azure
Rationale:
Configuring sampling and retention policies reduces the telemetry ingestion volume and costs, while preserving sufficient data to analyze trends in token usage and latency. Disabling tracing removes critical observability, increasing the context size increases token and cost pressure, and changing deployment tiers does not address data ingestion monitoring.
Share This Question
Challenge a friend or share with your study group.