JAS
← All insights
· 8 min readAI ResearchMIT StudyAgency StrategyAI Quality

MIT Tested AI on 11,500 Real Jobs. It Scored a 7 Out of 10. Here Is Why "Good Enough" Is Not Good Enough for Your Clients.

MIT FutureTech researchers tested 40+ AI models on 11,500 workplace tasks. The result: AI produces "minimally sufficient" work about 65% of the time. What that means for agencies selling quality.

Every CEO announcing AI-driven layoffs makes the same implicit claim: AI can do the job. MIT just tested that claim on 11,500 real workplace tasks. The answer is more nuanced, and more damaging to the "replace everyone" narrative, than either side wants to admit.

On 2 April 2026, MIT FutureTech researchers published one of the most comprehensive studies ever conducted on AI's actual workplace capabilities. They analysed 11,500 tasks from the U.S. Labor Department's database, tested them across more than 40 AI models using workplace-style prompts, and had 17,000+ AI-generated outputs evaluated by workers who actually do those jobs.

The headline finding: AI can complete roughly 65% of text-based tasks at what the researchers define as a "minimally acceptable" level. That is a 7 out of 10. Not excellent. Not reliable. Minimally sufficient.

What "Minimally Sufficient" Actually Means

The distinction between "can do the task" and "can do the task well" is the entire story. MIT did not find that AI produces excellent work 65% of the time. They found it produces work that clears the lowest bar of acceptability about two-thirds of the time.

In 2024, AI could complete approximately 50% of text-based tasks at this minimally acceptable level. By 2025, that rose to roughly 65%. The researchers project that by 2029, AI will handle 80 to 95% of text-based tasks, but still only at the "good enough" threshold.

The key finding that gets lost in the headlines: high-quality, error-free work remains significantly harder for AI to produce. The gap between "minimally sufficient" and "reliably excellent" is where the real story lives.

Consider what a 7 out of 10 means in practice. A 7 out of 10 job advert might attract candidates but miss crucial qualifying criteria that waste the hiring manager's time. A 7 out of 10 ad campaign might generate clicks but fail to communicate the brand's actual value proposition. A 7 out of 10 candidate assessment might surface relevant skills but miss the cultural fit indicators that determine whether the hire succeeds.

Minimally sufficient is not the standard any client pays premium rates for.

Rising Tide, Not Crashing Wave

The MIT researchers deliberately titled their paper "Crashing Waves vs. Rising Tides" because their data contradicts the dominant narrative on both sides of the AI employment debate.

The "AI will replace everyone" camp is wrong. The impact is gradual, not sudden. There is no mass wipeout coming. AI capabilities are improving steadily, not in dramatic leaps that render entire professions obsolete overnight.

The "AI is just hype" camp is also wrong. The capabilities are real and improving. Moving from 50% to 65% task coverage in a single year is significant progress. The trajectory points toward broad capability, eventually.

The reality, according to MIT, is that AI's impact will be broad (affecting many jobs across many industries) but gradual (playing out over years, not months). This is important for agencies because it means the disruption is not a single event to survive but a continuous shift to navigate.

The Gap Between Theory and Practice

One of the most underreported findings in the MIT study is the gap between theoretical AI capability and actual deployment. While models may technically be able to attempt 65% of tasks, the percentage being successfully deployed in real workplaces is significantly lower.

This mirrors findings from other recent research. Anthropic, the company that builds the Claude AI system, published its own study in March 2026 finding "limited evidence that AI has affected employment to date." While theoretical models suggest 94% of tasks in computer and mathematics occupations are "exposed" to AI, actual AI coverage sits at around 33%.

The gap between what AI can theoretically do and what it is actually doing in practice is massive. Companies are firing people based on theoretical capability, not proven deployment.

What This Means for Agency Clients

When a client replaces their agency with AI tools, they are making a specific bet: that "minimally sufficient" output is good enough for their business objectives. For some tasks, it is. For the tasks that drive revenue, build brand equity, and create competitive advantage, it is not.

The MIT data gives agencies a powerful counter-narrative. When a client says "AI can do what you do," the response is: "AI can attempt what we do. It produces minimally acceptable results about two-thirds of the time. Would you accept a 7 out of 10 on your most important campaigns?"

This is not an anti-AI argument. It is a quality argument. AI is a tool that can handle execution at an acceptable level for routine tasks. Strategy, judgment, quality assurance, and the ability to produce work that exceeds "minimally sufficient": that is what agencies charge for.

The Verification Tax

Related research adds another dimension to the MIT findings. Employees currently spend an average of 4.3 hours per week verifying AI-generated outputs. At typical salary levels, this represents approximately $14,200 per employee per year in verification overhead.

When companies fire their agency and bring work in-house using AI tools, they do not just get the AI output. They get the verification burden. Someone internal must check every piece of AI-generated content, every report, every analysis for errors, hallucinations, and quality issues.

This hidden cost rarely appears in the business case for replacing an agency. The projected savings assume AI output is usable as-is. The reality is that someone must still do the quality control, and that someone is now an internal employee who was not hired for that purpose.

LLM hallucinations alone cost businesses $67.4 billion in 2024. When 47% of business executives admit to making major decisions based on unverified AI-generated content, the cost of "minimally sufficient" becomes clearer.

The Agency Opportunity

The MIT study does not argue against using AI. It argues for using AI properly: as a tool that augments human judgment rather than replacing it.

For agencies, this creates a clear positioning opportunity. The agencies that will win in the next three to five years are not the ones that refuse to use AI. They are the ones that use AI to handle the 65% of tasks it can manage at an acceptable level, freeing their human team to focus on the 35% that requires expertise, judgment, and quality that exceeds the minimum.

This is the Publicis model. Publicis did not replace its people with AI. It used AI to make its people's output more valuable. The result: record profit margins and 5,800 new hires while every competitor was cutting headcount.

The agencies that position themselves as "AI plus human expertise" rather than "AI instead of humans" are aligned with what MIT's data actually shows. The ones that try to compete with AI on the tasks AI can already do at a minimally sufficient level will find themselves in a race to the bottom.

AI scores a 7 out of 10. Your agency should be selling the difference between 7 and 10, and making that difference measurable enough that clients understand exactly what they lose when they settle for minimally sufficient.