Why Africa’s Agriculture AI Must Listen Before It Speaks
By Najeeb G. Abdulhamid¹, Elizabeth “Liz” Ankrah¹, Abiodun Ogunyemi², Mark Perry³, Merja Lina M. Bauters², Jona Repishti⁴, Steven Sam³, Samuel Chege Maina¹, Millicent Ochieng¹, Mercy Muchai¹,...
View ArticleFrom Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning...
Wenyue Hua, Zachary Huang, Tyler Payne, Safoora Yousefi, Saleema Amershi, Asli Celikyilmaz Increasingly, we ask AI agents to act for us: to schedule our meetings, compare offers, settle the terms of a...
View ArticleTreatment Effect Assessment at Scale: Accounting for Correlated Metrics and...
By Kai Qi and Momo Meng, Microsoft Experimentation Platform At Microsoft’s Experimentation Platform (ExP), feature launches are routinely evaluated through A/B tests that can generate hundreds—or even...
View ArticleTemporal Augmentations for Streamed Video Games: Supplementary Material
This supplementary website accompanies the paper “Augmentations for Robust and Efficient Imitation Learning in Streamed Video Games,” published at the Conference on Games 2026. The paper studies...
View ArticleAnnouncing the AI Economy Institute’s Third Cohort of Senior Fellows
Continuing to Build a Global, Multidisciplinary Community of Knowledge The 2026 cohort of the AI Economy Institute (AIEI) brings together researchers from across North America, Europe, the Middle...
View ArticleStudy: Developing tangibles with and for people living with chronic illness
The study is a part of a larger goal of improving quality of life for people living with chronic illnesses through the use of small, tangible devices. About this study The purpose of this study is to...
View ArticleObject-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement
By Kinam Kim, Namiko Saito, Heecheol Kim, Katsushi Ikeuchi, Jaegul Choo and Yasuyuki Matsushita Video playback requires cookie consent Introduction A residual RL policy learns purely in simulation...
View ArticleSentinelBench, a Benchmark for Long-Running Monitoring Agents
By Matheus Kunzler Maldaner, Adam Fourney, Amanda Swearngin, Hussein Mozannar,Gagan Bansal, Maya Murad, Rafah Hosn, Saleema Amershi Video playback requires cookie consent Modern AI agents are...
View ArticleVITRA Redefines VLA Pre-training Paradigms via Human Video Reconstruction
When you see robots participating in running races or performing folk dances on stage, you might envision a future where a simple natural language command is all it takes for a robot to tidy up a...
View ArticleFara1.5 – A family of frontier computer use agent models
By: Ahmed Awadallah, Sahil Gupta, Yash Lara, Yadong Lu, Hussein Mozannar, Akshay Nambi, Zach Nussbaum, Yash Pandya, Aravind Rajeswaran, Corby Rosset, Alexey Taymanov, Luiz do Valle, Vibhav Vineet,...
View ArticleWhimsical Strategies Break AI Agents: Generating Out-of-Distribution...
By Zachary Huang, Tyler Payne, Gagan Bansal, Will Epperson, Wenyue Hua, Adam Fourney, Amanda Swearngin, Maya Murad, Ece Kamar, Saleema Amershi As AI agents are increasingly deployed to handle real...
View ArticleWebwright: A Terminal Is All You Need For Web Agents
Webwright GitHub repo Webwright project page By Yadong Lu1, Lingrui Xu2, Chao Huang2, Ahmed Awadallah11Microsoft Research and 2The University of Hong Kong Instead of solving web tasks by predicting...
View ArticleEvaluating Proactive AI Mediators in Multi-Party Conversation with ProMediate
By Ziyi Liu (opens in new tab), Bahar Sarrafzadeh, Pei Zhou, Longqi Yang (opens in new tab), Ashish Sharma Imagine you are in a high-stakes group discussion, stuck in a circular argument with no...
View ArticleThe Art of Building Verifiers for Computer Use Agents
By Corby Rosset, Pratyusha Sharma, Andrew Zhao, Miguel Gonzalez-Fernandez, Ahmed Awadallah We share lessons learned from building a best-in-class verifier for computer use agent trajectories on the...
View ArticleMemento: Teaching LLMs to Manage Their Own Context
Vasilis Kontonis, Yuchen Zeng, Shivam Garg, Lingjiao Chen, Hao Tang, Ziyan Wang, Ahmed Awadallah, Eric Horvitz, John Langford, Dimitris Papailiopoulos We taught models to compress their own...
View ArticleActions Speak Louder Than Prompts: Rethinking How LLMs Reason Over Graph Data
By Ben Finkelshtein (opens in new tab) (University of Oxford), Silviu Cucerzan, Sujay Kumar Jauhar, and Ryen W. White (Microsoft) Think about the last time you opened a shared document at work....
View ArticleExperiential Reinforcement Learning
By Taiwei Shi, Sihao Chen, Longqi Yang, Jaime Teevan Reinforcement Learning is at the core of building and improving frontier AI models and products. Yet most state-of-the-art RL methods learn...
View ArticleFrom One to Many
By Jaime Teevan, Chief Scientist & Technical Fellow In recent years we’ve all lived through the transition to cloud computing, a sudden shift to remote work, and now the rapid rise of AI. Each...
View ArticlePhi-Ground: Improving how AI agents navigate screen interfaces
Imagine an AI assistant that can navigate a computer the same way humans do—clicking buttons, filling out forms, and moving between applications—all by simply interpreting what’s on the screen. This...
View ArticleDeep Video Discovery: Using agentic search to analyze long-form video
Extracting useful information from long videos, whether meeting recordings, experimental data, or lecture content, requires painstaking manual review. AI tools offer some help: language-vision models...
View Article