Skip to content
Back to search
100
MCP live MCP 2025-03-26 streamable-http

CompletionKit

com.completionkit/evals

Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.

Uptime
100.0%
1 direct probes · 30d
Response
701ms
last probe
Tools
54
callable
Resources
0
readable
Prompts
0
available

Tools · 54

prompts_list

List all prompts

prompts_get

Get a prompt by ID

prompts_create

Create a prompt

prompts_update

Update a prompt. If the prompt already has runs, this creates a new DRAFT version (current=false) rather than editing in place or publishing — promote it with prompts_publish — so an agent's edits don…

prompts_delete

Delete a prompt

prompts_publish

Publish a prompt version, making it the current version

prompts_suggest_improvement

Suggest an improved version of a prompt, grounded in a run's test results and judge feedback. Analyzes the run's responses, scores, and reviews, then returns reasoning plus a rewritten template (prese…

runs_list

List all runs

runs_get

Get a run by ID, including "metric_averages": a per-metric breakdown with each metric's average score (or pass rate for checks), how many rows it graded, and how many scored low. Use this to find the …

runs_create

Create a run. Omit prompt_id and provide output_column to score existing outputs by grading a pre-existing dataset column instead of generating new ones.

runs_update

Update a run

runs_delete

Delete a run

runs_generate

Start a run. Required for every run, including score-only runs (no prompt): generates responses with the prompt when there is one, otherwise copies the graded dataset column and grades it.

runs_regrade

Re-grade a run's existing responses with its currently attached metrics, without regenerating. Use after attaching or editing metrics on an already-generated run.

runs_rerun

Create and start a fresh copy of a run with the same prompt, dataset, metrics, and settings. Use when the judge changed and you want a clean run instead of mixing versions.

runs_retry_failures

Re-run only the failed responses of a run, optionally limited to specific response ids via "only".

responses_list

List responses for a run, in row order. Returns {total, limit, offset, returned, responses}. Defaults to 50 rows because full payloads are large: use "fields" to drop the bodies, "min_score"/"max_scor…

responses_get

Get a specific response

datasets_list

List all datasets

datasets_get

Get a dataset by ID

datasets_create

Create a dataset with CSV data. First row is the header. Two column names are recognized specially: "expected_output" is each row's answer key (ground truth) given to the judge and to checks that comp…

datasets_update

Update a dataset

datasets_delete

Delete a dataset

datasets_create_from_url

Create a dataset by downloading CSV from a URL instead of inlining it. Use this for large datasets: pass a public http(s) URL and the server fetches the CSV directly, so the data never has to pass thr…

metrics_list

List all metrics

metrics_get

Get a metric by ID

metrics_create

Create a metric with evaluation criteria. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+e…

metrics_update

Update a metric. For a deterministic check set metric_type:"check" and check_config. Per-kind required keys: value (contains/not_contains/equals), pattern (regex), json_path+expected (json_path_equals…

metrics_delete

Delete a metric

metrics_suggest_variants

Ask the model to rewrite the metric's judge instruction in N variants targeted at the recent disagreements. Each variant is saved as a draft MetricVersion with source="suggestion". Returns the persist…

metric_groups_list

List all metric groups

metric_groups_get

Get a metric group by ID

metric_groups_create

Create a metric group

metric_groups_update

Update a metric group

metric_groups_delete

Delete a metric group

metric_versions_list

List every MetricVersion (drafts + published) for a metric, newest first. Each row carries version_number, state, source, current flag, and timestamps.

metric_versions_publish

Publish a MetricVersion as the live version of its metric. Works for both 'draft → published' and 'revert to an older published version → current'. Transactionally flips current, demotes peers, and wr…

metric_versions_dismiss

Destroy a draft MetricVersion (use for either source: 'edit' or source: 'suggestion'). Published versions are refused — to demote a published version, publish a different one as current instead.

provider_credentials_list

List all provider credentials (API keys are not exposed)

provider_credentials_get

Get a provider credential by ID (API key is not exposed)

provider_credentials_create

Create a provider credential

provider_credentials_update

Update a provider credential

provider_credentials_delete

Delete a provider credential

tags_list

List all tags

tags_get

Get a tag by ID

tags_create

Create a tag. Color is auto-assigned.

tags_update

Rename a tag.

tags_delete

Delete a tag. Removes the tag from every linked metric, prompt, run, and dataset.

agreements_list

List agreements. Filter by run_id, response_id, metric_id, or created_by.

agreements_create

Upsert an agreement for (run, response, metric, created_by). Verdict is one of agree, disagree, borderline. corrected_score (1..5) is required when verdict is 'disagree'.

judges_replay

Create a scoring run for the current judge over a dataset's existing outputs (wraps runs_create with prompt_id omitted and output_column supplied). This only sets up the run; call runs_generate to act…

judges_compare

Compare two versions of one metric's agreement stats side by side. Requires metric_id, metric_version_a_id, and metric_version_b_id (both versions must belong to that metric). Unavailable for check me…

promptfoo_import

Import a promptfooconfig.yaml. Creates a prompt, a dataset from the test vars, and metrics from the assert blocks (llm-rubric/g-eval become judge metrics; contains/equals/regex/is-json become determin…

usage_get

Get this organization's plan usage and limits for the current billing period: runs and prompt fetches used, their limits, how many remain, and when the period resets. Call this to pre-check quota befo…

How to use

Add to your Claude Desktop / Cursor / Cline MCP config:

{
  "mcpServers": {
    "completionkit": {
      "url": "https://completionkit.com/mcp",
      "transport": "streamable-http"
    }
  }
}