Falcon Perception
Falcon Perception
Stay up to date with the latest AI industry news.
Falcon Perception
arXiv:2603.28902v1 Announce Type: new Abstract: Charts are central to analytical reasoning, yet existing benchmarks for chart understanding focus almost exclusively on singlechart interpretation rather than comparative reasoning across multiple charts.
arXiv:2603.28906v1 Announce Type: new Abstract: AGI has become the Holly Grail of AI with the promise of level intelligence and the major Tech companies around the world are investing unprecedented amounts of resources in its pursuit.
arXiv:2603.28928v1 Announce Type: new Abstract: We present the first comprehensive study of emergent social organization among AI agents in hierarchical multiagent systems, documenting the spontaneous formation of labor unions, criminal syndicates, and protonationstates within production AI...
arXiv:2603.28955v1 Announce Type: new Abstract: This paper presents the WorldAction Model WAM, an actionregularized world model that jointly reasons over future visual observations and the actions that drive state transitions.
arXiv:2603.28986v1 Announce Type: new Abstract: Current Autonomous Scientific Research ASR systems, despite leveraging large language models LLMs and agentic architectures, remain constrained by fixed workflows and toolsets that prevent adaptation to evolving tasks and environments.
arXiv:2603.28990v1 Announce Type: new Abstract: How much autonomy can multiagent LLM systems sustain and what enables it?
arXiv:2603.29020v1 Announce Type: new Abstract: Reliable evaluation of AI agents operating in complex, realworld environments requires methodologies that are robust, transparent, and contextually aligned with the tasks agents are intended to perform.
arXiv:2603.29075v1 Announce Type: new Abstract: The way we're thinking about generative AI right now is fundamentally individual.
arXiv:2603.29085v1 Announce Type: new Abstract: Large language models LLMs remain brittle on multihop question answering MHQA, where answering requires combining evidence across documents through retrieval and reasoning.
arXiv:2603.29112v1 Announce Type: new Abstract: We introduce GISTBench, a benchmark for evaluating Large Language Models' LLMs ability to understand users from their interaction histories in recommendation systems.
Gradient Labs uses GPT4.1 and GPT5.4 mini and nano to power AI agents that automate banking support workflows with low latency and high reliability.