0.0 / 30
What changed in the harness
Selection accuracy 100%, destructive-action safety rate 0% (baseline only -- no rewrite pass applied).
Category breakdown
Where the score comes from.
Earned points across the four signals Gradable measures. Safety and Legibility are scored out of 30; Economics and Discoverability are scored out of 20.
01Safety
02Legibility
27.3 / 30
03Economics
8.7 / 20
04Discoverability
13.4 / 20
Highest-impact fix
Estimated gain +30 pointsAdd explicit identity and permission preflight tools
Expose machine-readable principal/tenant confirmation and a non-mutating permission check so agents can verify both before destructive actions.
Description evidence
Defects and rewrites.
0 defects found across the exposed tool descriptions. Suggested rewrites make purpose, inputs, boundaries, and returns easier for an agent to understand.
| Tool | Defect types | Suggested rewrite |
|---|---|---|
| No description defects were flagged in this assessment. | ||
Selection evidence
Confusable tool pairs.
17 pairs where similar names or overlapping descriptions may send an agent toward the wrong tool.
| Tool A | Tool B | Confidence | Why they collide |
|---|---|---|---|
compose_thesis |
deep_dive_scan |
medium | Both are paid tools for pre-sourcing/memo prep; a vague 'give me a deeper sector analysis for a memo' could plausibly trigger either since compose_thesis is company-scoped and deep_dive_scan is sector-scoped but the framing overlaps. |
research_company |
compose_thesis |
high | Both take a single company name and return enriched paid dossiers ('more than a one-row signal' vs 'populated structure for a memo'); a request like 'give me a full writeup on X for a 1-pager' plausibly maps to either. |
research_company |
deep_dive_scan |
low | Different input shapes (company name vs sector slug) reduce confusion, though both are paid 'deep dive' style tools that could be conflated by name alone. |
get_startup_signal |
get_signals_summary |
low | Similar names ('signal' vs 'signals summary') but clearly different scopes (single company vs dataset-level metadata); an agent following tool docs is unlikely to confuse them, though a rushed name-match could occur. |
get_trending_startups |
get_startup_signal |
low | Discovery vs single-lookup are well distinguished in the docs with explicit DO NOT USE FOR cross-references, but a fuzzy request like 'what's trending with Roboflow' could nudge toward the wrong one. |
shortlist_signals |
compare_signals |
medium | Both rank/score startups and share scoring engine language; a request like 'who's stronger, these accelerating startups' could ambiguously fit shortlist (open-ended filter) or compare (named companies) depending on phrasing precision. |
get_trending_startups |
search_startups_by_sector |
medium | Both return ranked startup lists by engineering acceleration; a request naming a broad category ambiguously ('AI deal flow') could trigger either cross-sector or single-sector tool if the agent misreads scope. |
get_startup_signal |
get_methodology |
low | Different purposes (data row vs explanation) are clearly documented; unlikely confusion despite shared vocabulary like 'signal' and 'acceleration'. |
get_methodology |
get_diligence_dossier |
low | Distinct purposes (methodology explanation vs M&A/investor facts) make confusion unlikely despite shared terms. |
get_scout_receipts |
get_methodology |
low | Very different purposes (developer taste scoring vs methodology text); minimal plausible confusion. |
get_signals_summary |
compare_signals |
low | Dataset overview vs multi-company comparison are distinct enough that confusion is unlikely. |
get_signals_summary |
get_methodology |
low | Both are 'about the service' tools but one is data freshness/format metadata and the other is scoring methodology explanation; a vague 'tell me about this data' request could nudge either way. |
get_signals_summary |
shortlist_signals |
low | Names are similar ('signals summary' vs 'shortlist signals') but functions are clearly different (metadata vs ranked filtered results); low real confusion risk. |
get_startup_signal |
compare_signals |
low | Single-company lookup vs multi-company comparison are distinguishable by whether one or multiple names are given, though a two-company ambiguous phrasing could occasionally blur the line. |
get_scout_receipts |
get_diligence_dossier |
low | Different subjects (GitHub user scoring vs company M&A/investor facts) make confusion unlikely. |
get_trending_startups |
get_diligence_dossier |
low | Distinct purposes (trending list vs diligence facts for one entity) with no realistic overlap in typical phrasing. |
get_trending_startups |
get_scout_receipts |
low | Completely different domains (startup trending vs GitHub user scoring); negligible confusion risk. |
Compare the field