Skip to content

Illustrative scenario · Synthetic data · No client engagement or measured business outcome

OFFER H1 · SCN-02 · PARTIAL

Can a product assistant pass a release review?

An imaginary B2B software company has a document-answering prototype. The demonstration uses a newly authored fictional product manual, synthetic user accounts, conflicting document versions, and deliberately unanswerable questions.

FIG. 01 · EVALUATION PATH · SCN-02Evaluation setFROM YOUR OWN INPUTSFeature under testUNCHANGEDScore vs rubricPER INPUT CLASSPassCOUNTED, NOT ASSUMEDFailure taxonomyRANKED BY CONSEQUENCEGo / no-go memoCONDITIONS NAMEDCORRECTFAILED
The feature is not modified. It is scored against a rubric on inputs drawn from your own data, and the failures are ranked by consequence rather than counted. The memo names the conditions that would change the decision.

WHAT WE WOULD SHOW

  • A question answered with cited evidence
  • A refusal when the evidence is missing
  • Two accounts with different permissions getting correctly different answers
  • A version conflict between two revisions of the manual
  • An escalation to a human reviewer
  • An evaluation report that includes the failures, not only the passes

PROPOSED ARTIFACTS

  • Synthetic test questions
  • Expected-answer rubrics
  • Evaluation results
  • A release decision memo

ILLUSTRATIVE TEST DESIGN, NOT RESULTS

Use 100 independently reviewed synthetic tasks as a starting test set. Example targets could be at least 90% rubric pass rate and zero access violations in a defined permission test suite.

Those are candidate acceptance criteria for a fictional exercise. They are not a demonstrated capability or a general safety guarantee. Observed results would be published only after running the demo, alongside the exact dataset version, method, model configuration, limitations, and failed cases.

Have a feature with unclear release criteria? Discuss the evaluation problem.