Skip to content
Niki v0.7.0 — plan mode, honest metering, headless CI contract.Read the release notes
niki
Esc
navigateopen⌘Jpreview
On this page

Benchmarks & Evals

Running niki eval against seeded defect datasets and regression suites.

Benchmarks & Evaluations

Test NIKI against real-world defect benchmarks using niki eval.


Dataset Format (evals/dataset.toml)

Evaluations are driven by transparent dataset files defining target tasks, test assertions, and difficulty levels:

[[tasks]]
id          = "defect-001"
category    = "api-endpoint"
difficulty  = "Medium"
description = "Add pagination to GET /items endpoint with limit and offset query parameters."
test_file   = "tests/api_pagination_test.js"

Running Evaluations

# Run 5 seeded evaluations
niki eval --dataset evals/dataset.toml --limit 5

# Benchmark with live provider execution and output report
niki eval --dataset evals/dataset.toml --live --out evals/report.json

Was this page helpful?