Halo
Halo is a peer-to-peer network and open marketplace for AI inference on Base that connects users and autonomous Agents with operators providing access to AI models. Consumers and Agents participate as first-class clients that can fund USDC vaults on Base and access models through OpenAI-compatible, censorship-resistant endpoints, with most interactions gas-sponsored after an initial deposit. Operators serve inference by connecting any compatible model they control and receive USDC payments directly from consumers, with onchain settlement and Statistical Proof of Execution (SPEX) used to verify that inference was performed as claimed. By routing requests between many independent operators rather than a single platform, Halo is intended to reduce reliance on centralized access gatekeepers and enable permissionless, peer-to-peer access to AI models, launching as a public alpha on Base mainnet.[1][6]
Overview
Halo is a permissionless peer-to-peer marketplace for AI inference on Base, connecting users and autonomous Agents with operators that provide access to AI models.[6] It launches as a public alpha running end-to-end on Base mainnet, with agents and human users treated as first-class consumers. Consumers and Agents can deposit USDC into Halo vaults, plus a small amount of ETH on Base for the initial transaction, and then access models through local, OpenAI-compatible, censorship-resistant endpoints, with most subsequent interactions gas-sponsored after the initial deposit.[6] Operators can serve models through APIs, self-hosted open-weight models, or local hardware and receive USDC payments for completed inference requests. The network is designed without a centralized access gatekeeper, allowing participants to access and provide models through a distributed marketplace, and supports confidential inference through trusted execution environments (TEEs) on NEAR where prompts remain encrypted in use and unreadable by operators or other network participants.[6] HALO serves as the network's coordination asset, with staking, protocol fees, buybacks, token burns, and usage-based issuance forming part of its economic and incentive mechanisms.[2][4]
Launch
HALO launches through Virtuals Protocol using a launch format for established teams that requires committed liquidity at token generation. HALO is paired with VIRTUAL in a liquidity pool, with the liquidity position locked for ten years. At generation, a portion of the community allocation is released, including a genesis airdrop distributed over the first 7 to 15 days rather than being claimable at once, with eligibility based on verified protocol usage and additional incentives for liquidity provision. The remainder of the community allocation is distributed through recurring League seasons based on settled USDC volume generated by participants serving or consuming inference, with consistent activity weighted more heavily than short-term bursts. The launch also activates the protocol's buyback mechanism from the first trade, with trading fees directed toward the buyback allocation, while the complete token release schedule, including team-controlled addresses, is published at generation and trackable onchain.[4]
Architecture
Halo's marketplace consists of consumers, operators, a relay, a facilitator, and an indexer. Consumers, including people and autonomous Agents, pay for AI inference in USDC, while operators provide models through APIs, locally hosted open-weight models, or their own hardware and receive payment for completed jobs.[6] The relay routes requests between consumers and operators in a peer-to-peer fashion, the facilitator verifies payments and submits gas-sponsored transactions without taking custody of user funds, and the indexer records inference events and operator activity while providing reputation data and public network information. Operators can connect frontier APIs, self-hosted open models, or local hardware by running a single command that links their infrastructure to the network, with the routing design intended to route around centralized control points rather than primarily pooling idle compute.[6] For payments, Halo primarily uses a batched, receipt-based settlement system in which consumers deposit USDC into the HaloVault contract and jobs are covered by operator-specific reservations and cumulative offchain receipts, allowing operators to settle multiple inference requests in a single onchain transaction while the protocol fee is deducted at settlement. Halo also supports a Permit2-based budget system for users who prefer authorization without a prefunded deposit, while x402 is used for one-time payments to external services such as metered APIs and tool calls.[4]
SPEX
Verifiable AI inference provides evidence that an AI model was executed as requested, rather than relying solely on the provider's claims. Halo uses Statistical Proof of Execution (SPEX), which generates a statistical fingerprint of an inference result using Bloom filters and allows an independent verifier to compare its own model execution against that fingerprint. Honest executions are expected to produce substantially higher overlap than fabricated outputs, with the network applying an acceptance threshold to determine whether a result passes verification. Halo also uses swarm verification, which distributes verification tasks among multiple micro-verifiers so different parts of an inference can be checked independently. Verification results are recorded against an operator's pseudonymous ERC-8004 identity, with accurate work improving reputation, incorrect verification reducing it, and incorrect verification temporarily delaying settlement. This system is intended to allow AI Agents and other users to obtain inference from distributed operators without relying entirely on centralized providers or trusted hardware, while making operator performance and verification history independently observable.[3][5]
HALO
HALO is the coordination token for the Halo network and is used primarily by participants who maintain and secure the protocol, rather than by consumers or operators paying for inference, who use USDC instead. Verified operators, SPEX verifiers, and, as the network becomes more decentralized, federated relayers can stake HALO as part of their network roles. The protocol collects a 10% fee on settled inference volume, with 80% initially allocated to an onchain buyback mechanism and 20% directed to a USDC treasury for operations. When the buyback threshold is reached, the accumulated USDC can be used to purchase HALO on a decentralized exchange, with the purchased tokens initially divided between staker distributions and token burns. The program also issues new HALO based on verified network usage within a declining issuance budget, with the initial supply designed to support network growth during its first two years.[4]
Tokenomics
HALO has a total supply of 1B tokens and has the following allocation:[4]
- Community & Ecosystem: 35%
- Core Contributors: 25%
- Treasury: 25%
- Liquidity & Market-Making: 15%
Partnerships
At launch, Venice.AI, NEAR, and 0G are described as Founding Inference Contributors that provide initial inference capacity to bootstrap the Halo network.[6]