Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
arXiv:2604.20987v1 Announce Type: new Abstract: Long horizon interactive environments are a testbed for evaluating agents skill usage abilities.
AI 업계의 최신 소식을 빠르게 확인하세요.
arXiv:2604.20987v1 Announce Type: new Abstract: Long horizon interactive environments are a testbed for evaluating agents skill usage abilities.
arXiv:2604.20995v1 Announce Type: new Abstract: Alignment faking, where a model behaves aligned with developer policy when monitored but reverts to its own preferences when unobserved, is a concerning yet poorly understood phenomenon, in part because current diagnostic tools remain limited.
arXiv:2604.21003v1 Announce Type: new Abstract: AI agents are increasingly deployed on complex, domainspecific workflows navigating enterprise web applications that require dozens of clicks and form fills, orchestrating multistep research pipelines that span search, extraction, and synthesis,...
arXiv:2604.21006v1 Announce Type: new Abstract: We introduce Deep FinResearch Bench, a practical and comprehensive evaluation framework for deep research DR agents in financial investment research.
arXiv:2604.21018v1 Announce Type: new Abstract: While scaling testtime compute can substantially improve model performance, existing approaches either rely on static compute allocation or sample from fixed generation distributions.
arXiv:2604.21027v1 Announce Type: new Abstract: Electronic health record EHR question answering is often handled by LLMbased pipelines that are costly to deploy and do not explicitly leverage the hierarchical structure of clinical data.
arXiv:2604.21036v1 Announce Type: new Abstract: TexttoimageT2I models like Stable Diffusion and DALLE have made generative AI widely accessible, yet recent studies reveal that these systems often replicate societal biases, particularly in how they depict demographic groups across professions.
arXiv:2604.21044v1 Announce Type: new Abstract: In some complex domains, certain problemspecific decompositions can provide advantages over monolithic designs by enabling comprehension and specification of the design.
On Friday, Chinese AI firm DeepSeek released a preview of V4, its longawaited new flagship model. Notably, the model can process much longer prompts than its last generation, thanks to a new design that helps it handle large amounts of text more efficiently.
Meta has been poaching talent from Thinking Machines Lab. But it's a twoway street.
새만금개발청청장 직무대리 정인권은 새만금을 미래형 첨단 인공지능AI 스마트도시로 본격 조성하기 위해 법정계획인 ‘새만금 스마트도시계획2026~2030년’을 국토교통부의 최종 승인을 거쳐 24일, 공고하였다고 밝혔다.이번 계획은 새만금을 탄소중립과 AI 기반의 미래 혁신 도시로 조성하기 위한 중장기 전략으로, 글로벌기업 협업과 첨단기술을 기반으로 한 선도형 스마트도시 모델을 제시한 것이 특징이다.이를 실현하기 위해 ‘탄소중립 AI혁신 스마트도시 새만금’을 비전으로 설정하고, 친환경 에너지 및 탄소중립 기반 생태계 구축, A
ComfyUI, whose tools give creators more control over AI image, video, and audio generation, just raised $30 million.