Skip to content

Engineering Blueprint

Safety checked

PromptFoo LLM Evaluation Pipelines

3Vaibhav S.Machine Learning EngineerG2September 2026

Ensure LLM applications behave correctly by building deterministic evaluation pipelines that gate deployments on real-world test cases and semantic quality metrics. PromptFoo configs let you test actual application behavior with model-graded rubrics and versioned datasets across your CI workflow.

1 File Included

  • eval-builder-skill.md

    24 KB

What problem does this solve?

Build or extend language-agnostic PromptFoo evaluation pipelines for LLM apps and agents, including providers, datasets, rubrics, deterministic gates, and GitHub Actions CI. Use for eval design or integration work, not generic unit tests or telemetry-only observability.

How does it work?

  1. Inspect the target codebase to inventory existing eval configs, providers, datasets, rubrics, and CI workflows.
  2. Define the output contract and curate or assess the evaluation dataset against the application's real use cases, happy paths, and corner cases.
  3. Build a PromptFoo config with a provider that invokes production code, deterministic assertions for verifiable properties, and model-graded rubrics for semantic properties.
  4. Add a GitHub Actions workflow triggered on pull requests to run the eval matrix, upload results and reports, and optionally publish a summary comment. Output: a versioned, language-neutral evaluation pipeline that gates on real application behavior.

What's the biggest win?

Production prompts and agent behavior change automatically in eval tests without copying test-only fixtures or duplicating code.

What's required to run this?

  • PromptFoo CLI: installed via Node/npm/pnpm/yarn/bun package manager, matched to repository version and lockfile
  • Provider strategy depends on application runtime: import production code directly for JavaScript/TypeScript; invoke real Python entrypoint through a Python provider or exec for Python apps; use Ruby provider or exec for Rails; use exec or HTTP for Go, Java, and other runtimes
  • Deterministic assertions: JSON schema, required fields, enums, bounds, counts, ordering, forbidden strings, tool names, and tool-call sequence
  • Model-graded checks: llm-rubric, g-eval, or context-faithfulness for relevance, groundedness, tone, safety behavior, and quality; keep judge temperature fixed at 0
  • Fixture data in YAML format; rubric text in .txt files; file:// paths resolve relative to config directory for vars and relative to project root for assertions
  • Retry strategy: bounded retries for transient failures only (network, timeout, 408/409/429/5xx); do not retry authentication, invalid-request, parser, schema, deterministic, or rubric failures; use exponential backoff with jitter and cap attempts and wall-clock time
  • GitHub Actions workflow: trigger on pull_request events, use path filters to avoid spending API budget on unrelated changes, set timeout and concurrency cancellation, run with --no-cache for fresh calls, upload machine-readable JSON and HTML reports with if: always()
  • For G2 ai-playbooks repository: use evals/ layout with promptfoo/, fixtures/, and rubrics/ subdirectories; import production prompts directly from src/ so eval exercises the same code as runtime; use load-env.ts for environment setup; fixture tool responses and cap conversation turns to prevent state bleed across test cases
  • No PR-comment reporting implemented by default; artifact-only reporting is the baseline behavior in established repositories until explicitly requested

Tools in this Blueprint

ChatGPT logo
4.6(2,984 reviews)
PromptFoo
GitHub Actions
Node.js

About This Blueprint

Industry
Engineering
ProductivitySafety checked

Email Triage GSuite

Surface what actually needs you and draft replies in your voice, with memory that carries items forward so nothing slips.

AL

Aaron LeBlanc

CEO

ProductivitySafety checked

Email Triage Outlook

Stop missing commitments and get a daily brief that surfaces only what genuinely needs you.

AL

Aaron LeBlanc

CEO

EngineeringSafety checked

AI-Assisted Incident Root Cause Analysis

Rapidly synthesize fragmented incident evidence into a structured root-cause diagnosis, distinguishing correlation from causation and confidence levels. The workflow builds a validated timeline, evaluates competing hypotheses, and produces remediation steps without inventing conclusions when evidence remains incomplete.

JW

Jayesh Wankhede

Software Engineer II