When the Algorithm Writes the Copy: Where AI Agents in Marketing Actually Work — and Where They Don't
- Lucas Blochberger
- 5 days ago
- 9 min read
Klarna rolled out an AI assistant in February 2024 that handled 2.3 million conversations within a month — two-thirds of all customer service interactions. Resolution time: under two minutes instead of the previous eleven. Repeat inquiries down 25 percent. CEO Sebastian Siemiatkowski sold it as proof that AI could replace human labor at scale.
In May 2025, he walked it back. Speaking to Bloomberg, he said cost had been "a too predominant evaluation factor," that quality had suffered, and that genuine investment in human support was "the way of the future." Not a clean retreat, not a failure across the board. The volume numbers held up. What broke was quality at the edges — complex, multi-step, emotionally charged cases. Klarna now runs a hybrid model.
This one story contains almost everything you need to know about AI agents in marketing in 2026. They work reliably on narrowly scoped, measurable tasks. They break down on anything requiring judgment, nuance, or genuine accountability. And anyone who only reads the first press release gets half the truth.
Key Takeaways
AI agents work reliably on narrowly scoped, measurable tasks — ad targeting, first drafts, simple support triage — and fail at anything requiring judgment, nuance, or genuine accountability
Over 40% of agentic AI projects are expected to be canceled by end of 2027, according to Gartner — not because of weak models, but due to cost, unclear business value, and inadequate risk controls
88% of agent pilots never reach production, according to Forrester/Anaconda; MIT research puts the figure at 95% with no measurable P&L impact
AI SDRs are the year's most expensive disappointment: just 15% conversion to qualified opportunities versus 25% for human reps, plus elevated spam and reputation risk
Visibly AI-generated marketing content tends to cost trust rather than build it — transparency labeling makes evaluations more critical, not more favorable
Human-in-the-loop isn't a transitional stage until the models improve — it's currently the model that consistently wins

The Numbers Show a Gap, Not a Curve
86.4 percent of marketing teams use AI in at least one area, according to HubSpot's 2026 State of Marketing Report — up from 41 percent in 2024. Adoption has arrived. Success hasn't. Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027. Not because of weak models, but because of runaway costs, unclear business value, and inadequate risk controls. Forrester and Anaconda, in a joint study, put the figure at 88 percent of agent pilots that never reach production. The biggest blockers, in order: evaluation gaps, governance friction, unreliable models.
A widely cited MIT study (Project NANDA, summer 2025) claims that 95 percent of GenAI pilots showed no measurable P&L impact. The number has circulated in every other LinkedIn post since. It's based on 52 interviews and roughly 300 public deployments, using a narrow definition of success ("ROI within six months"). It's worth citing — but not selling as established fact. That's exactly what's happening industry-wide right now, and it's a small example of the hype this article is dissecting.
Where It Works
Ad targeting and budget allocation are the least glamorous case, but the most solid. Meta's Advantage+ delivers an average 22 percent higher ROAS, by the company's own account. That figure comes from the vendor, no question — but independent practitioner consensus at least confirms the direction: targeting automation works, creative automation doesn't yet. Experienced advertisers let Meta decide on bids and audiences while keeping creative control in-house.
Tier-1 customer service triage is the second solid case, with Klarna serving as both a positive and a cautionary example. Salesforce reports impressive Agentforce references: 1-800Accountant with 90 percent case deflection during tax season, Reddit with 46 percent deflection and response times that dropped from 8.9 to 1.4 minutes, OpenTable with 70 percent of inquiries resolved autonomously. All vendor numbers. Salesforce's own CRMArena-Pro benchmark measures only around 35 percent accuracy for an unconfigured Agentforce agent, and independent estimates suggest fewer than ten percent of customers have scaled past pilot status. Industry-wide median deflection for tier-1 support sits at around 41 percent in 2026, with the top quartile just under 59 percent. The strongest success factor is mundane: how many systems the agent can see live. Knowledge base only: around 28 percent deflection. Knowledge base plus CRM: around 38 percent. Skip the integration work, and you don't get the showcase numbers either.
First drafts for content are the third case, and the least contested. McKinsey puts the ROI for AI-assisted content drafting at 3.2x, and personalization at 2.7x. HubSpot measures around 6.1 hours saved per week per marketer. That tracks with what's now standard in every editorial team: AI delivers the first version, a human edits, trims, and corrects the tone. Useful as an assistant, not as an autopilot.
Where It Fails
AI SDRs — automated B2B outreach — are this year's most expensive disappointment. The market grew fast: the vendor 11x raised a $50 million Series B from a16z, and the market for AI SDR tools crossed the $4 billion mark in 2025. Independent results tell a different story. AI SDRs convert booked meetings to qualified opportunities at around 15 percent; human SDRs at around 25 percent. In head-to-head comparisons, human sales reps generated 2.6x the revenue and achieved significantly higher show rates. Add to that a reputation problem many underestimate: AI-generated emails land in spam at around eight percent, compared to around three percent for human-written ones, and domains running AI outbound at production volume lose roughly 38 reputation points within 90 days. The result: between 40 and 60 percent of AI SDR pilots die within three months, some with domain damage that can only be undone by buying a new, aged domain. Autonomous agents optimize for volume. Volume past a certain threshold destroys deliverability — exponentially, not linearly.
Content at scale is the second failure case, and Google made clear in 2026 how it evaluates that. The March 24, 2026 spam update — completed in under 20 hours, the fastest confirmed rollout ever — hit programmatic SEO sites with thin AI content hard: traffic losses between 60 and 80 percent, sometimes more. Google doesn't penalize AI itself, but what it internally calls "scaled content abuse" — mass output without editorial oversight. Sites that used AI as a tool within an editorial process were largely unaffected. Even prominent brands got hit. HubSpot's own blog reportedly lost 70 to 80 percent of its organic traffic after publishing at volumes far outside its core expertise.
Complex, emotional customer service is the third failure case. 75 percent of customers report that chatbots fail on complex issues. A message combining a medical context, a complaint, and a billing error reliably overwhelms most systems.
The Embarrassments That Stick
DPD had to shut down its chatbot in January 2024 after a customer got it to swear and write a poem calling DPD the "worst delivery firm in the world." The post reached 1.3 million views. Coca-Cola released AI-generated Christmas ads in 2024 and 2025, both dismissed by critics as soulless; according to CARMA sentiment analysis, approval dropped from 23.8 to 10.2 percent. McDonald's Netherlands pulled its AI-generated Christmas ad after more than 37,000 posts accusing it of "AI slop" appeared on the day of release alone. Deloitte Australia had to refund part of a government contract after an AI-assisted report contained fabricated sources that slipped through multiple review layers. A Canadian tribunal ordered Air Canada to pay compensation after its chatbot gave incorrect information about bereavement fares. The airline's defense — that the bot was a "separate legal entity" — was rejected.
The Trust Problem Nobody Solves
Only seven percent of consumers say visibly AI-generated marketing content strengthens their trust in a brand. 31 percent say the opposite. The share of people who trust a heavily AI-reliant brand less rose from 20 to 39 percent within a year. Over 80 percent want AI content clearly labeled as such. That's exactly where the trap lies, as a study by the Nuremberg Institute for Market Decisions revealed: merely labeling content as AI-generated makes evaluations more critical, not more trusting. Identical ads were rated as less natural and less useful the moment an AI label appeared next to them. Transparency doesn't build trust here — it costs engagement. For the European market, where labeling requirements are coming, that's not a footnote.
What This Means in Practice
Andrej Karpathy, AI researcher and OpenAI co-founder, put it this way in October 2025: "They just don't work. They don't have enough intelligence, they're not multimodal enough, they can't do computer use and all this stuff. They don't have continual learning… They're cognitively lacking and it's just not working." He calls the coming decade "the decade of agents," not the year. Gartner analyst Anushree Verma sees most agentic projects as "early-stage experiments or proof of concepts that are mostly driven by hype and are often misapplied." Of the thousands of vendors calling themselves "agentic," Gartner estimates only a few hundred actually live up to the claim.
The practical takeaway is unspectacular but solid. AI agents belong where tasks are narrow, rule-based, and backed by clean data: first drafts, reporting, ad targeting, simple support triage. They don't belong where final creative judgment, brand voice, complex sales conversations, or genuine emotional situations are required. Anyone rationalizing people out of those areas because a pilot showed good volume numbers is repeating Klarna's first mistake. Human-in-the-loop isn't a transitional stage until the models get better — it's simply the model that's currently winning, consistently. And anyone building AI visibly into customer-facing content should know it's more likely to cost trust than earn it.
FAQ: AI Agents in Marketing
Where do AI agents work most reliably in marketing?
On narrowly scoped, rule-based tasks backed by clean data: ad targeting, first drafts for content, reporting, and simple support triage. Meta's Advantage+ delivers an average 22% higher ROAS; McKinsey puts the ROI for AI-assisted content drafting at 3.2x.
Why do so many agentic AI projects fail?
Not because of weak models. Gartner cites runaway costs, unclear business value, and inadequate risk controls as the main reasons — over 40% of projects are expected to be canceled by end of 2027. Forrester and Anaconda estimate that 88% of agent pilots never reach production.
What's the lesson from the Klarna example?
Klarna's AI assistant took over two-thirds of customer service in 2024 with strong volume numbers, but broke down on complex, emotionally charged cases. In 2025, Klarna reinvested deliberately in human support — a hybrid model instead of full automation.
Do AI SDRs work in B2B sales?
Rarely reliably. AI SDRs convert booked meetings to qualified opportunities at only around 15%, compared to around 25% for human SDRs. AI-generated emails also land in spam significantly more often, putting domain reputation at risk.
Does visibly AI-generated marketing content hurt brand trust?
Generally, yes. Only 7% of consumers say visible AI labeling strengthens their trust; 31% say the opposite. Simply labeling content as AI-generated makes evaluations more critical rather than more trusting.
What does this mean for using AI agents in practice?
AI agents belong where tasks are narrow, rule-based, and backed by clean data. They don't replace final creative judgment, brand voice, or complex, emotional customer situations. Human-in-the-loop isn't currently a transitional stage — it's the model that consistently wins.
About the Author
Lucas Blochberger is the founder and CEO of Blck Alpaca OG, a Vienna-based agency for data-driven marketing and AI agent integration. With a three-person team, he builds AI systems for enterprise clients across the DACH region, from SEO automation to custom agent architecture. His view of AI in marketing is unsentimental. Coming from a technical background himself, he knows the gap between vendor numbers and production reality — and writes about what actually works, not what sells well.
More from Lucas: blckalpaca.at
Sources
Klarna: AI assistant handles two-thirds of customer service chats in its first month (Klarna, Feb. 2024)
Klarna's course change: Klarna flips from AI-first to hiring people again (Fortune, May 2025)
HubSpot: 2026 State of Marketing Report
Gartner: Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (June 2025)
MIT Project NANDA: MIT Report Finds 95% of AI Pilots Fail to Deliver ROI (Summary; Original report "The GenAI Divide: State of AI in Business 2025")
Salesforce: Better LLM Agents for CRM Tasks (CRMArena Pro Benchmark)
Salesforce: Agentforce 360 Launch – Customer Testimonials (Oct. 2025)
11x funding: 11x.ai raises $50M Series B led by A16Z (TechCrunch)
AI-SDR Conversion Rates: AI SDR vs. Human SDR: What the Data Says (Autobound, SuperAGI Analysis)
Google Spam Update: Google Begins Rolling Out The March 2026 Spam Update (Search Engine Journal)
DPD Chatbot: DPD Disables AI Chatbot After It Swears At Customer (Silicon UK, Jan. 2024)
McDonald's Ad: Not lovin' it: McDonald's pulls AI-generated Christmas ad (NBC News)
Air Canada: Air Canada found liability for chatbot's bad advice on bereavement rates (CBC News, Feb. 2024)
Andrej Karpathy: OpenAI cofounder Andrej Karpathy says it will take a decade before AI agents actually work (Business Insider via AOL, Oct. 2025; Original: Dwarkesh Podcast)