Engineering Blueprint
Safety checkedPromptFoo LLM Evaluation Pipelines
Ensure LLM applications behave correctly by building deterministic evaluation pipelines that gate deployments on real-world test cases and semantic quality metrics. PromptFoo configs let you test actual application behavior with model-graded rubrics and versioned datasets across your CI workflow.
1 File Included
eval-builder-skill.md
24 KB
What problem does this solve?
Build or extend language-agnostic PromptFoo evaluation pipelines for LLM apps and agents, including providers, datasets, rubrics, deterministic gates, and GitHub Actions CI. Use for eval design or integration work, not generic unit tests or telemetry-only observability.
How does it work?
- Inspect the target codebase to inventory existing eval configs, providers, datasets, rubrics, and CI workflows.
- Define the output contract and curate or assess the evaluation dataset against the application's real use cases, happy paths, and corner cases.
- Build a PromptFoo config with a provider that invokes production code, deterministic assertions for verifiable properties, and model-graded rubrics for semantic properties.
- Add a GitHub Actions workflow triggered on pull requests to run the eval matrix, upload results and reports, and optionally publish a summary comment. Output: a versioned, language-neutral evaluation pipeline that gates on real application behavior.
What's the biggest win?
Production prompts and agent behavior change automatically in eval tests without copying test-only fixtures or duplicating code.
What's required to run this?
- PromptFoo CLI: installed via Node/npm/pnpm/yarn/bun package manager, matched to repository version and lockfile
- Provider strategy depends on application runtime: import production code directly for JavaScript/TypeScript; invoke real Python entrypoint through a Python provider or exec for Python apps; use Ruby provider or exec for Rails; use exec or HTTP for Go, Java, and other runtimes
- Deterministic assertions: JSON schema, required fields, enums, bounds, counts, ordering, forbidden strings, tool names, and tool-call sequence
- Model-graded checks:
llm-rubric,g-eval, orcontext-faithfulnessfor relevance, groundedness, tone, safety behavior, and quality; keep judge temperature fixed at 0 - Fixture data in YAML format; rubric text in .txt files; file:// paths resolve relative to config directory for vars and relative to project root for assertions
- Retry strategy: bounded retries for transient failures only (network, timeout, 408/409/429/5xx); do not retry authentication, invalid-request, parser, schema, deterministic, or rubric failures; use exponential backoff with jitter and cap attempts and wall-clock time
- GitHub Actions workflow: trigger on pull_request events, use path filters to avoid spending API budget on unrelated changes, set timeout and concurrency cancellation, run with --no-cache for fresh calls, upload machine-readable JSON and HTML reports with if: always()
- For G2 ai-playbooks repository: use evals/ layout with promptfoo/, fixtures/, and rubrics/ subdirectories; import production prompts directly from src/ so eval exercises the same code as runtime; use load-env.ts for environment setup; fixture tool responses and cap conversation turns to prevent state bleed across test cases
- No PR-comment reporting implemented by default; artifact-only reporting is the baseline behavior in established repositories until explicitly requested
Tools in this Blueprint
About This Blueprint
- Industry
- Engineering
More Blueprints to explore
AI-Assisted Incident Root Cause Analysis
Rapidly synthesize fragmented incident evidence into a structured root-cause diagnosis, distinguishing correlation from causation and confidence levels. The workflow builds a validated timeline, evaluates competing hypotheses, and produces remediation steps without inventing conclusions when evidence remains incomplete.
Jayesh Wankhede
Software Engineer II