AI agents: what can they actually do?
From a good answer to a finished task. A practical look at the capabilities, the limits, and the human decisions in between.
Not just
about technology.
About people, too.
Independent
AI magazine
Since 2026
Explanations, practical guides, and a closer look at the research.
From a good answer to a finished task. A practical look at the capabilities, the limits, and the human decisions in between.
GitHub’s walkthrough shows an agent building a board around the task.
Choose the task, the hardware, and the operating boundaries before choosing the model.
Cohere researchers trace how cultural diversity narrows through training data.
Hume AI researchers examine the human quality of voice interaction.
IBM evaluates modernization across a project rather than an isolated function.
Google DeepMind describes layered oversight for increasingly capable internal agents.
OpenAI explains changes to the infrastructure behind real-time conversation.
Hugging Face examines Pro, Flash, and the needs of long-running agents.
Mistral’s Spaces engineering story considers interfaces for people and agents.
ServiceNow researchers look at both task completion and the human experience of voice AI.
Google DeepMind proposes a cognitive framework for evaluating AI progress.
Mistral engineers describe a specialized loop for generating and checking tests.
Anthropic publishes a revised account of how it wants the model to behave.
Hugging Face looks beyond one breakout model to a changing community.
Anthropic’s Economic Index adds a more detailed view of how people use Claude.
A new architecture paper is a reminder that bigger is only one way a model can change.
Start the year with a better question: what, exactly, did the test measure?