TECHNICAL SPECIFICATION • V0.1.0

Transparent Reliability Methodology

AgentProof measures autonomous onchain agents using deterministic, reproducible formulas and SSRF-hardened network probes. Every metric displayed in our Passports and API traces directly back to the mechanisms documented below.

1. Identity vs. Runtime Operability

Onchain standards such as ERC-8004 prove that an agent identity exists and was minted on BNB Chain. However, onchain registration says nothing about whether the advertised service, API endpoint, or agent-to-agent protocol is currently reachable or operational.

Core Principle: AgentProof treats onchain registration exclusively as an identity declaration, never as evidence of operational uptime. Runtime reliability is continuously and independently measured.

2. The 5 Deterministic Probe Types

Autonomous probes execute against agent services on a scheduled cadence using our hardened probe engine:

01METADATA_RESOLUTION

Resolves the agent's advertised metadata URI over HTTP(S) or IPFS gateways, enforcing strict 1MB response size limits.

02SERVICE_REACHABILITY

Executes DNS resolution and TCP/TLS handshake with DNS pinning to verify that the host is reachable from public IP networks.

03HTTP_STATUS

Checks that the endpoint returns standard valid HTTP response codes (2xx/3xx/405/422). 5xx errors or connection drops record as failure.

04RESPONSE_LATENCY

Measures high-resolution round-trip time in milliseconds (median and P95 percentiles) from probe dispatch to header arrival.

05PROTOCOL_RESPONSE_VALIDITY

Validates that the returned payload adheres to expected MIME types and JSON structure without malformed syntax.

3. Availability & Latency Formulas

Measured Availability is computed over explicit sliding windows (24 Hours, 7 Days, and 30 Days):

Measured Availability % = ( Successful Attributable Probes / Total Attributable Probes ) * 100

Attributable vs. Excluded Outcomes

To guarantee rigorous fairness, AgentProof partitions probe outcomes into two strict categories:

✓ Attributable Outcomes (Count in Denominator)
  • SUCCESS: Endpoint responded within parameters.
  • AGENT_UNREACHABLE: Agent server dropped connection or refused TCP.
  • DNS_FAILURE: Agent hostname failed public resolution.
  • TIMEOUT: Agent service exceeded the 10-second timeout.
  • PROTOCOL_INVALID: Agent returned HTTP 5xx or malformed payload.
⊘ Excluded Outcomes (Never Penalize Availability)
  • BLOCKED_BY_SECURITY_POLICY: Target was an RFC1918/localhost IP blocked by runner security policy.
  • UPSTREAM_INDEXER_FAILURE: 8004scan or RPC gateway was unavailable.
  • AGENTPROOF_INTERNAL_ERROR: Runner internal execution error.

4. Evidence Sufficiency Tiers

Every reliability percentage is paired with an explicit sufficiency tier that communicates sample depth:

STRONG EVIDENCE30+ observations spanning at least 75% of the window duration. Statistically robust sample.
MODERATE EVIDENCE10–29 observations with regular temporal spread. Representative operational profile.
LIMITED EVIDENCE3–9 observations. Early measurement history; displayed with preliminary sample notice.
INSUFFICIENT EVIDENCEFewer than 3 observations. Availability percentage is intentionally withheld to prevent unrepresentative conclusions.

5. Onchain Reputation Evidence & Reviewer Distribution

AgentProof analyzes 8004scan onchain feedback data to evaluate reviewer diversity and concentration using non-accusatory statistical metrics:

  • Herfindahl-Hirschman Reviewer Concentration (HHI): Measures whether feedback is dominated by a small number of wallet addresses.
  • Reviewer Diversity Ratio: The fraction of total feedback records submitted by unique wallets (uniqueReviewers / totalFeedback).
  • Neutral Signal Taxonomy: Signals such as LOW_REVIEWER_DIVERSITY or HIGH_REVIEWER_CONCENTRATION describe empirical distribution shapes without subjective accusations or blacklisting.

6. Security & Probe Policy

Probing arbitrary third-party endpoints carries inherent SSRF risks. AgentProof implements strict security controls tested by 36 adversarial IP-policy tests:

DNS PinningResolves IP once and connects strictly to verified public IP addresses, preventing DNS rebinding attacks.
RFC1918 & Cloud BlocklistImmediately terminates probes targeting 10.x, 172.16.x, 192.168.x, 127.0.0.1, or 169.254.169.254 (AWS/GCP metadata).
Ethical Rate LimitingGlobal concurrency cap (10), per-host cap (2), minimum 5-second interval, and automatic cooldown backoff.

7. Explicit Limitations & Boundaries

  • Reachability is not correctness: A successful HTTP 200 response proves the agent's server is online, not that its internal AI reasoning or financial transactions are bug-free.
  • Append-only application level: Observations are append-only at the application layer. Proofs are not yet committed to zero-knowledge rollups.
  • No financial guarantees: AgentProof evidence is an informational metric for builders, orchestrators, and indexers.