Samples
Realistic Jev decision templates with labelled evaluation data.
Support ticket triage
Speculative fan-out classifies category, urgency, and escalation in one Jev call.
4 labelled test cases
LLM guardrail jailbreak detection
Detects prompt injection attempts and routes by attack pattern before an agent sees the text.
4 labelled test cases
Content moderation
Mild, non-graphic moderation labels plus a severity score for triage queues.
4 labelled test cases
RAG passage relevance
Scores how directly a retrieved passage answers a user query.
4 labelled test cases
Citation and claim verification
Checks whether a source supports, contradicts, or lacks enough information for a claim.
4 labelled test cases
Intent routing for agent tools
Maps user requests to a tool or no-op before invoking an agent workflow.
4 labelled test cases
Sales lead qualification
Composite BANT-style scoring: average budget, authority, need, and timeline scores.
4 labelled test cases
Product review sentiment
Scores sentiment from 1 to 5 and picks the dominant product aspect.
4 labelled test cases
Entity alignment and dedup
Uses object instructions with backtick references to decide if records represent the same company.
4 labelled test cases
Pull-request risk assessment
Scores change risk and flags whether a security review is warranted from a diff summary.
4 labelled test cases
Email phishing detection
Detects likely phishing and labels the social-engineering tactic.
4 labelled test cases
Confidence-gated refund routing
Demonstrates threshold gating: high-confidence refunds auto-accept, ambiguous cases go to review.
4 labelled test cases