Public leaderboard

Public assessment

8randonpickart5/alderpost-mcp (alderpost-mcp)

alderpost-mcp · v1.2.2 · scanned

What changed in the harness

Selection accuracy 95→95, token cost up 2%, unconfirmed writes 0%→0%.

Category breakdown

Where the score comes from.

Earned points across the four signals Gradable measures. Safety and Legibility are scored out of 30; Economics and Discoverability are scored out of 20.

01Safety

0.0 / 30

0.0 out of 30
02Legibility

26.8 / 30

26.8 out of 30
03Economics

20.0 / 20

20.0 out of 20
04Discoverability

17.5 / 20

17.5 out of 20

Highest-impact fix

Estimated gain +30 points

Add explicit identity and permission preflight tools

Expose machine-readable principal/tenant confirmation and a non-mutating permission check so agents can verify both before destructive actions.

Description evidence

Defects and rewrites.

8 defects found across the exposed tool descriptions. Suggested rewrites make purpose, inputs, boundaries, and returns easier for an agent to understand.

Tool Defect types Suggested rewrite
domain_shield
no_return_description
Domain security scan with VirusTotal malware detection (70+ antivirus engines). SPF, DKIM, DMARC, SSL, MX, DNSSEC, domain info. Returns a report of 8 checks scored 0-100 with prioritized recommendations. Price: $0.12 USDC on Base.
company_xray
no_return_description
Company intelligence powered by People Data Labs + Hunter.io. Returns a scored profile covering industry, employee count, revenue estimate, tech stack, verified email contacts, and social presence across 9 data sources scored 0-100. Price: $0.15 USDC on Base.
threat_pulse
no_return_description
Threat intelligence with VirusTotal (70+ engines) + AbuseIPDB (30K+ community reporters). Returns a scored threat report covering blacklists, reverse DNS, open ports, SSL analysis, and email security across 7 sources scored 0-100. Price: $0.10 USDC on Base.
compliance_check
no_return_description
IT security compliance audit with Qualys SSL Labs grading. Returns letter grades for 8 checks covering email auth, SSL/TLS, OWASP headers, cookies, privacy, DNSSEC, hosting, and data exposure. Price: $0.15 USDC on Base.
prospect_iq
no_return_description
Sales intelligence powered by People Data Labs + Hunter.io. Returns a scored prospect profile covering web presence, tech stack, verified contacts, social signals, and contact readiness across 7 sources scored 0-100. Price: $0.12 USDC on Base.
sports_edge
no_return_description
Pre-game sports intelligence with ESPN + The Odds API + Claude AI synthesis. Returns live standings, odds comparisons from 15+ bookmakers, and AI-generated game analysis. Price: $0.12 USDC on Base.
property_intel
no_return_description
Location intelligence with US Census Bureau + OpenWeather data. Returns a scored report covering geocoding, nearby amenities, schools, elevation, demographics, and walkability across 7 sources scored 0-100. Price: $0.10 USDC on Base.
health_signal
no_return_description
Health intelligence with NIH RxNorm drug interaction checking + FDA data. Returns a risk-scored report covering drug labels, adverse events, recalls, nutrition, and drug-to-drug interactions with severity ratings across 7 sources scored 0-100. Price: $0.10 USDC on Base.

Selection evidence

Confusable tool pairs.

1 pair where similar names or overlapping descriptions may send an agent toward the wrong tool.

Tool A Tool B Confidence Why they collide
company_xray prospect_iq medium Both are near-identical intelligence tools (same providers, same domain input, both rate sources 0-100) and overlap on tech stack and verified contacts. A task like 'get an intelligence report on this company's tech stack and contacts' maps to both; only firmographics (employees/revenue) vs. sales signals (contact readiness) separates them, so an agent may pick the wrong one.

Compare the field

One score is useful.
The evidence makes it actionable.

Back to the leaderboard