Neuroscientists, military personnel, and even a prisoner: this is how the team that 'hacks' Microsoft's AI before it reaches the public wo

The company has a "red team" that evaluates all artificial intelligence systems before their launch, and halts them if necessary.
From left to right, Daniel Kluttz, Ram Shankar, Siva Kumar, and Tori Westerhoff at Microsoft headquarters in Redmond, Ohio (USA). MICROSOFT
Seattle - March 20, 2026 -
Brad Smith, president of Microsoft, pauses for a moment to reflect and uses the word "guardrails" with the nonchalance of someone who has given a lot of thought to precipices. A conference on the company's innovation is being held at its headquarters in Redmond, Washington, to which this and other international newspapers have been invited, and EL PAÍS asks him how and by whom it is determined whether the company's artificial intelligence (AI) can be used in a war context, such as the current one . Just a few days ago, it became public that the artificial intelligence firm Anthropic has sued the Pentagon for banning it after drawing red lines on the use of its technology. This is the current debate in the world of Big Tech, and it is a very familiar matter for Microsoft: in 2021, the Pentagon canceled a $10 billion agreement with the company after protests from its employees. Microsoft, in fact, has supported Anthropic in its fight with the Pentagon.
Smith responds: “We have principles, we define them, and we publish them. By definition, those principles create guardrails. And we stay on the road within them. It’s not just about when we should use technology, but also when we shouldn’t.”
Microsoft has a team dedicated to hacking its own products: the "red team ." The name has military roots. Red teams originated in militaries to simulate enemy attacks and detect their own vulnerabilities before the real adversary could. In cybersecurity, the practice has been established for decades. But applying it to generative artificial intelligence is relatively new, and Microsoft claims to have pioneered it by forming this team in 2018. "Before a product is released, red teams break the technology so others can rebuild it stronger and more secure," explains Ram Shankar Siva Kumar , a self-described "data cowboy " and leader of the red team. "AI can cause problems, from security breaches to psychosocial harm. People use Copilot [la IA de Microsoft] in moments of high vulnerability, so observing how these systems can fail before they reach the user is crucial," he explains.
This sort of internal AI task force has already analyzed more than 100 of the company's products. Microsoft doesn't disclose how many people work on it, nor whether or which products have been paused. But it does assert that the team has the authority to do so: “No high-risk AI system is deployed without first undergoing independent testing. If our team identifies serious risks that haven't been mitigated, the product isn't released until those issues are resolved,” Kumar says.
The question the team asks itself when analyzing a product before it is launched is: "How could this AI system be used, for good or for bad, months or years from now?"
Six principles
The “guardrails” Smith mentioned are six overarching principles that the team believes are very clear when examining products: fairness, accountability, transparency, reliability and security, inclusion, and privacy and security. These principles translate into concrete tools on a daily basis. “If you give an engineer a fifty-page document to implement those principles, they’re going to be overwhelmed. We have an open-source tool called Pyrit ; we built it for ourselves and then made it available to the world because we believe in a healthy ecosystem,” says Kumar.
The red team includes neuroscientists, linguists, national security specialists, cybersecurity experts, military veterans, and even a former prisoner "who was rehabilitated," Kumar explains. They also speak 17 languages, and "some dialects of French, Mongolian, Thai, and Korean," according to the team leader, since one of the red team 's obsessions, he explains, is ensuring that the AI doesn't make mistakes anywhere in the world.
Alongside Kumar, Tori Westerhoff leads the operations of the red team. Her background combines cognitive neuroscience—she studied at Yale and was one of the first members of the Wharton Neuroscience Initiative —with national security strategy, having worked in intelligence and defense agencies. “When we receive an assignment,” she explains, “we emulate what could go wrong at the extremes of that technology’s usage curve. My team delves into how to use that product as intended, and in unintended ways, to identify the most extreme cases and help the product team reproduce and mitigate them before they can be used by anyone in the real world.”
One example of their work was the "red teaming ," as they internally call their hacking , of GPT-5, the OpenAI (Microsoft partner) model launched last August. What they did was train another AI to attempt to hack the program, automatically and on a scale impossible for humans.
When testing GPT-5, the red team used Pyrit to automatically generate more than two million trap conversations . The attacking AI tried to deceive the targeted AI, relentlessly, for days, exploring combinations a human would never think of. Finding these weaknesses manually is an incredibly slow process; therefore, they trained this AI to try to break another AI, “like in Inception ,” says Kumar, referring to the Christopher Nolan film where the characters enter dreams within dreams.
However, Westerhoff, Kumar, and Daniel Krutz, who heads the company's Responsible AI office, emphasize one point: automation has its limits. " Red teaming can only be automated to a certain extent, and only humans can determine whether an AI-generated response makes them uncomfortable or represents bias," the company asserts. The criteria are set by the individual; the scale is determined by the machine. This division of labor defines the team's philosophy.
Westerhoff believes that, in fact, only the human mind is capable of “imagining those spaces that have not yet been observed, that have not been fully defined or explored; our job is to innovate and create beyond the space that has been systematized.”
The team identifies three areas where automation is inherently blind and human judgment is essential. The first concerns subject matter; people are needed to assess risk in areas such as medicine or security. The second relates to the environments where AI is deployed; “ we need humans to account for linguistic differences and redefine what constitutes harm in different political and cultural contexts,” the company says. And the third is emotional intelligence. Ultimately, only humans can assess the range of interactions users might have with AI systems. A model can pass all automated tests and still produce responses that are disturbing to a real person in a specific situation.
This view of AI aligns with that of Mustafa Suleyman , one of the founders of DeepMind (now part of Google) and CEO of Microsoft AI. A few days ago, he wrote in the journal Nature : “A seemingly sentient AI can be weaponized.” As artificial intelligence systems increasingly mimic the structure of human language, he argues, we need design rules and laws to prevent them from being mistaken for sentient beings. “They must remain fundamentally accountable to humans and subordinate to the well-being of humanity,” Suleyman writes. “AI agents should have no more rights or freedoms than my laptop.”
The central philosophy that guides the work of the red team is, ultimately, that “responsible AI is not a filter applied at the end of development, but a foundational part of the process,” says Kumar. They are like Smith's guardrails, which don't actually act as brakes, but rather as a condition for going fast without going over the edge.
El Pais, Spain






