Abstract:
With the rapid advancement of large language model (LLM)-driven AI agents, tool use has become the dominant paradigm through which agents dynamically interact with external services during task execution. In multi-tenant agent platforms, each agent sandbox often needs to access a wide range of heterogeneous external services concurrently, including LLM inference APIs, web scraping APIs, code space APIs, and financial data APIs. These APIs often require distinct network egress paths. Existing solutions each have significant limitations: tunnel-based approaches require establishing a separate channel per service, leading to maintenance costs that grow sharply with the number of services; IP-based policy routing struggles with the dynamically changing endpoints of API providers; and user-space proxies such as HAProxy and Envoy incur context-switching overhead and require intrusive configuration. To address these challenges, this paper proposes ASTRA, a transparent, high-performance name-based routing system. ASTRA leverages the existing TLS Server Name Indication (SNI) extension as a service identifier, enabling transparent routing compatible with the existing ecosystem without requiring modifications to applications or infrastructure. Meanwhile, it implements domain-name-based routing decisions entirely within the Linux kernel, preventing the context-switching overhead of user-space proxies to achieve high-performance forwarding. ASTRA incorporates three key techniques: (1) a cross-layer name-based routing scheme using the combination of TLS server name and TCP port as the service identifier; (2) a 4-way TCP-TLS joint handshake proxy model that transparently intercepts connections and extracts service names in kernel space; and (3) a deferred in-kernel name resolution mechanism based on a custom radix tree for efficient name-to-address translation. Experimental results show that, under typical API payload sizes, ASTRA outperforms HAProxy and Envoy in request rate by 27.16% and 50.74%, and in P50 latency by 17.68% and 43.60%, respectively. ASTRA maintains near-constant lookup latency as name rules scale from 1 000 to 50 000, showing scalability as the number of name rules grows. In terms of CPU resource efficiency measured by request rate per core, ASTRA achieves 1.91× the HAProxy baseline and 1.68× the Envoy baseline.