Transparent Reliability Methodology
AgentProof measures autonomous onchain agents using deterministic, reproducible formulas and SSRF-hardened network probes. Every metric displayed in our Passports and API traces directly back to the mechanisms documented below.
1. Identity vs. Runtime Operability
Onchain standards such as ERC-8004 prove that an agent identity exists and was minted on BNB Chain. However, onchain registration says nothing about whether the advertised service, API endpoint, or agent-to-agent protocol is currently reachable or operational.
2. The 5 Deterministic Probe Types
Autonomous probes execute against agent services on a scheduled cadence using our hardened probe engine:
Resolves the agent's advertised metadata URI over HTTP(S) or IPFS gateways, enforcing strict 1MB response size limits.
Executes DNS resolution and TCP/TLS handshake with DNS pinning to verify that the host is reachable from public IP networks.
Checks that the endpoint returns standard valid HTTP response codes (2xx/3xx/405/422). 5xx errors or connection drops record as failure.
Measures high-resolution round-trip time in milliseconds (median and P95 percentiles) from probe dispatch to header arrival.
Validates that the returned payload adheres to expected MIME types and JSON structure without malformed syntax.
3. Availability & Latency Formulas
Measured Availability is computed over explicit sliding windows (24 Hours, 7 Days, and 30 Days):
Attributable vs. Excluded Outcomes
To guarantee rigorous fairness, AgentProof partitions probe outcomes into two strict categories:
SUCCESS: Endpoint responded within parameters.AGENT_UNREACHABLE: Agent server dropped connection or refused TCP.DNS_FAILURE: Agent hostname failed public resolution.TIMEOUT: Agent service exceeded the 10-second timeout.PROTOCOL_INVALID: Agent returned HTTP 5xx or malformed payload.
BLOCKED_BY_SECURITY_POLICY: Target was an RFC1918/localhost IP blocked by runner security policy.UPSTREAM_INDEXER_FAILURE: 8004scan or RPC gateway was unavailable.AGENTPROOF_INTERNAL_ERROR: Runner internal execution error.
4. Evidence Sufficiency Tiers
Every reliability percentage is paired with an explicit sufficiency tier that communicates sample depth:
5. Onchain Reputation Evidence & Reviewer Distribution
AgentProof analyzes 8004scan onchain feedback data to evaluate reviewer diversity and concentration using non-accusatory statistical metrics:
- Herfindahl-Hirschman Reviewer Concentration (HHI): Measures whether feedback is dominated by a small number of wallet addresses.
- Reviewer Diversity Ratio: The fraction of total feedback records submitted by unique wallets (
uniqueReviewers / totalFeedback). - Neutral Signal Taxonomy: Signals such as
LOW_REVIEWER_DIVERSITYorHIGH_REVIEWER_CONCENTRATIONdescribe empirical distribution shapes without subjective accusations or blacklisting.
6. Security & Probe Policy
Probing arbitrary third-party endpoints carries inherent SSRF risks. AgentProof implements strict security controls tested by 36 adversarial IP-policy tests:
7. Explicit Limitations & Boundaries
- Reachability is not correctness: A successful HTTP 200 response proves the agent's server is online, not that its internal AI reasoning or financial transactions are bug-free.
- Append-only application level: Observations are append-only at the application layer. Proofs are not yet committed to zero-knowledge rollups.
- No financial guarantees: AgentProof evidence is an informational metric for builders, orchestrators, and indexers.