Q
A company uses a Microsoft Copilot Studio agent across multiple Dynamics 365 apps to generate natural language responses. The agent supports English and Spanish and relies on Dynamics 365 data and business logic.
The company regularly deploys updates to its Dynamics 365 environment by using CI/CD.
The company's development team wants to detect regressions in agent output quality caused by application updates. The company has anonymized production prompt logs available.
You need to recommend a testing approach to evaluate output quality at scale and detect regressions.
What should you recommend?
Question Info
Choose the Best Option
Click any option to instantly check if you're correct.
Explanation
Objective:
3.2 Manage the testing of AI-powered business solutions
What This Item Tests:
Validate effective Copilot prompt best practices
Additional Reading:
Options for regression testing
Rationale:
Automated batch testing using production prompt logs enables evaluation at scale with realistic inputs, supports multilingual validation, and allows consistent regression comparison across updates; Synthetic prompts and safety checks alone do not reflect real-world usage or measure overall quality; Manual review does not scale and cannot reliably detect regressions across large datasets; Prompt-level unit tests are not suitable because LLM outputs are non-deterministic and vary across runs.
Share This Question
Challenge a friend or share with your study group.