연구
Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory
arXiv:2608.20397v1 Announce Type: new Abstract: Agentic large language models LLMs on the Model Context Protocol MCP reencode verbose tool schemas every turn, so prefill quadratic in sequence length dominates timetofirsttoken TTFT as the tool registry grows.
이 콘텐츠는 ArXiv AI 원본 기사의 요약입니다. 전문은 원본 사이트에서 확인해주세요.
원문 기사 보기 →