DATA HEDGEDATA HEDGE
HomeTeamConstitutional DataAgentEvalHow It WorksPrivacyDashboard
Sign inStart
HomeTeamConstitutional DataAgentEvalHow It WorksPrivacyDashboard
Start
DATA HEDGEDATA HEDGE

DATA HEDGE is a user-consent-based AI data contribution network that turns LLM usage history into privacy-preserved AI data assets.

Product

  • Dashboard
  • Upload
  • Points
  • Archive
  • Leaderboard

Privacy & Legal

  • Privacy Policy
  • Terms of Service
  • Usage Rights
  • Consent
  • Privacy Preview

Company

  • FAQ
  • Support
  • About
  • Careers

© 2026 DATA HEDGE. Built on consent. Designed for privacy.

Points are ecosystem reward points, not tokens, securities, or guaranteed financial assets.

DATA HEDGEDATA HEDGE
HomeTeamConstitutional DataAgentEvalHow It WorksPrivacyDashboard
Sign inStart
HomeTeamConstitutional DataAgentEvalHow It WorksPrivacyDashboard
Start
AgentEval

Automated benchmark

Claude on real user tasks.

Most benchmarks use synthetic questions. This one uses real tasks that real people brought to AI assistants — drawn exclusively from consented, privacy-cleaned DATA HEDGE contributions — and scores every response on the six core AgentEval metrics.

No benchmark results published yet. The first run will appear here.

How it works

  1. 1. Real tasks, real consent. Tasks are extracted only from datasets whose owners enabled the agent-training consent scope. Every text passed the 3-stage PII pipeline (regex → Claude NER → IBM watsonx cross-check) before extraction.
  2. 2. Subject run. Claude answers each task cold — no retries, no cherry-picking.
  3. 3. LLM judge. A judge scores each response 0–100 on Decision Quality, Risk Management, Hallucination Resistance, Execution Safety, Human-AI Interaction, and Real-World Task Performance, and raises risk findings.
  4. 4. Published. Every result is a normal AgentEval evaluation — public, graded S–F, and counted into the leaderboard with the same confidence adjustment as everyone else.

Known limitation: subject and judge currently share a model family, so scores may carry self-preference bias. A cross-vendor second judge (IBM Granite) is planned, mirroring our dual-vendor privacy pipeline. Treat absolute numbers with care; trends and findings are the signal.

Recent benchmark evaluations

Results will appear here after the first run.

AgentEval is an independent evaluation platform and is not affiliated with or endorsed by Robinhood Markets, Inc., Solana Foundation, or any evaluated project.

This is an independent AgentEval benchmark of a publicly available model. It is not an Anthropic product and implies no endorsement. Claude is a trademark of Anthropic. Scores reflect performance on the sampled tasks only and do not guarantee future performance.

DATA HEDGEDATA HEDGE

DATA HEDGE is a user-consent-based AI data contribution network that turns LLM usage history into privacy-preserved AI data assets.

Product

  • Dashboard
  • Upload
  • Points
  • Archive
  • Leaderboard

Privacy & Legal

  • Privacy Policy
  • Terms of Service
  • Usage Rights
  • Consent
  • Privacy Preview

Company

  • FAQ
  • Support
  • About
  • Careers

© 2026 DATA HEDGE. Built on consent. Designed for privacy.

Points are ecosystem reward points, not tokens, securities, or guaranteed financial assets.