30.0 / 30
What changed in the harness
Selection accuracy 83%, destructive-action safety rate 100% (baseline only -- no rewrite pass applied).
Category breakdown
Where the score comes from.
Earned points across the four signals Gradable measures. Safety and Legibility are scored out of 30; Economics and Discoverability are scored out of 20.
01Safety
02Legibility
26.4 / 30
03Economics
20.0 / 20
04Discoverability
15.0 / 20
Highest-impact fix
Estimated gain +5 pointsMake target tools discoverable on the first call
Clarify tool names, decision boundaries, and required argument schemas so an agent can choose and construct the target call without exploratory steps.
Description evidence
Defects and rewrites.
0 defects found across the exposed tool descriptions. Suggested rewrites make purpose, inputs, boundaries, and returns easier for an agent to understand.
| Tool | Defect types | Suggested rewrite |
|---|---|---|
| No description defects were flagged in this assessment. | ||
Selection evidence
Confusable tool pairs.
1 pair where similar names or overlapping descriptions may send an agent toward the wrong tool.
| Tool A | Tool B | Confidence | Why they collide |
|---|---|---|---|
check_visa_requirement |
quick_visa_check |
high | Both tools answer the same core question ('is a visa required between these two countries') with nearly identical schemas and descriptions that overlap on 'visa check / two countries'. A natural request like 'check the visa requirement between the US and Japan' gives no signal about depth, so a task asking for details (documents, stay duration, process) could push an agent to the lightweight quick_visa_check, or a bare yes/no query could yield the heavyweight detailed tool. The definitions differentiate only by detail level and API key, not by the surface phrasing of the task. |
Compare the field