Last updated:
Halo is a peer-to-peer marketplace for AI inference that connects users and autonomous Agents with operators providing access to AI models. Its infrastructure includes onchain payments, model-serving interfaces, and Statistical Proof of Execution (SPEX) for verifying that inference was performed as claimed. [1]
Halo is a permissionless peer-to-peer marketplace for AI inference on Base, connecting users and autonomous Agents with operators that provide access to AI models. Consumers and Agents can deposit USDC into Halo vaults and use compatible models through an OpenAI-compatible endpoint, while operators can serve models through APIs, self-hosted open-weight models, or local hardware and receive USDC payments for completed inference requests. The network is designed without a centralized access gatekeeper, allowing participants to access and provide models through a distributed marketplace, and supports confidential inference through trusted execution environments (TEEs) on NEAR for workloads requiring additional privacy. HALO serves as the network's coordination asset, with staking, protocol fees, buybacks, token burns, and usage-based issuance forming part of its economic and incentive mechanisms. [2] [4]
HALO launches through Virtuals Protocol using a launch format for established teams that requires committed liquidity at token generation. HALO is paired with VIRTUAL in a liquidity pool, with the liquidity position locked for ten years. At generation, a portion of the community allocation is released, including a genesis airdrop distributed over the first 7 to 15 days rather than being claimable at once, with eligibility based on verified protocol usage and additional incentives for liquidity provision. The remainder of the community allocation is distributed through recurring League seasons based on settled USDC volume generated by participants serving or consuming inference, with consistent activity weighted more heavily than short-term bursts. The launch also activates the protocol's buyback mechanism from the first trade, with trading fees directed toward the buyback allocation, while the complete token release schedule, including team-controlled addresses, is published at generation and trackable onchain. [4]
Halo's marketplace consists of consumers, operators, a relay, a facilitator, and an indexer. Consumers, including people and autonomous Agents, pay for AI inference in USDC, while operators provide models through APIs, locally hosted open-weight models, or their own hardware and receive payment for completed jobs. The relay routes requests between consumers and operators, the facilitator verifies payments and submits gas-sponsored transactions without taking custody of user funds, and the indexer records inference events and operator activity while providing reputation data and public network information. For payments, Halo primarily uses a batched, receipt-based settlement system in which consumers deposit USDC into the HaloVault contract and jobs are covered by operator-specific reservations and cumulative offchain receipts, allowing operators to settle multiple inference requests in a single onchain transaction while the protocol fee is deducted at settlement. Halo also supports a Permit2-based budget system for users who prefer authorization without a prefunded deposit, while x402 is used for one-time payments to external services such as metered APIs and tool calls. [4]
Verifiable AI inference provides evidence that an AI model was executed as requested, rather than relying solely on the provider's claims. Halo uses Statistical Proof of Execution (SPEX), which generates a statistical fingerprint of an inference result using Bloom filters and allows an independent verifier to compare its own model execution against that fingerprint. Honest executions are expected to produce substantially higher overlap than fabricated outputs, with the network applying an acceptance threshold to determine whether a result passes verification. Halo also uses swarm verification, which distributes verification tasks among multiple micro-verifiers so different parts of an inference can be checked independently. Verification results are recorded against an operator's pseudonymous ERC-8004 identity, with accurate work improving reputation, incorrect verification reducing it, and incorrect verification temporarily delaying settlement. This system is intended to allow AI Agents and other users to obtain inference from distributed operators without relying entirely on centralized providers or trusted hardware, while making operator performance and verification history independently observable. [3] [5]
HALO is the coordination token for the Halo network and is used primarily by participants who maintain and secure the protocol, rather than by consumers or operators paying for inference, who use USDC instead. Verified operators, SPEX verifiers, and, as the network becomes more decentralized, federated relayers can stake HALO as part of their network roles. The protocol collects a 10% fee on settled inference volume, with 80% initially allocated to an onchain buyback mechanism and 20% directed to a USDC treasury for operations. When the buyback threshold is reached, the accumulated USDC can be used to purchase HALO on a decentralized exchange, with the purchased tokens initially divided between staker distributions and token burns. The program also issues new HALO based on verified network usage within a declining issuance budget, with the initial supply designed to support network growth during its first two years. [4]
HALO has a total supply of 1B tokens and has the following allocation: [4]