Skip to content
The Playbook

Amazon AI Agents: What Five Operators Actually Run in 2026

Updated August 31, 2026

Key Takeaways

  • An agent's IQ is not the problem, its context is. Ask a PPC agent about a 55% ACoS and it tells you to cut bids. Tell it the same ASIN has a 60% subscribe and save rate and a $42 first year customer value, and it tells you to raise the budget. The data did not change, only what the agent was told about the business.
  • Almost nobody has handed over real work yet. Across 450 SellerMate clients and 4,914 agent questions, only 23 clients have anything running on a schedule. Those 23 now drive one in five of every interaction the agent has.
  • Scheduled work finishes, typed questions wander. Scheduled agent tasks came back with a finished deliverable 95% of the time. Ad hoc chat questions finished 41% of the time. Same agent, same data.
  • The agent inherits whatever your listing data already gets wrong. Amazon's AVEN reads your photos as well as your text, so a listing that says V neck over a crew neck photo loses relevance before any agent touches it.
  • Amazon just split your title in two. A 75 character title plus 125 character item highlights, rolling out through the rest of 2026, with correction or search suppression for listings that ignore it. Across 12,546 executed edits, 52% of runs ended in no change at all.

I spoke on a live panel in August 2026 called Amazon Ops Agents: The Future of Seller Operations, hosted by Akash Singh at SellerMate. Four other operators presented alongside me, and the useful thing about the session was that none of us had compared notes beforehand. We all arrived at roughly the same conclusion from four completely different angles.

If you want the wider view of where this is heading, we covered it in how AI is changing Amazon selling. This piece is narrower: what five operators have actually put into production.

The conclusion we all landed on: the agent is rarely the thing that failed. What it was given to work with is where the problem started, and that is the part you control.

Here is what each speaker actually said, with the numbers they put on screen.

Panel lineup for the SellerMate Amazon Ops Agents webinar on 26 August 2026: Andrew Morgans of Marknology, Jon Tilley of ZonGuru, Monte Desai of Pixii, Ben Mathew of Superfuel AI, hosted by Akash Singh of SellerMate
The panel, 26 August 2026. Hosted by Akash Singh of SellerMate, with Jon Tilley of ZonGuru, Monte Desai of Pixii, Ben Mathew of Superfuel AI, and me.

Why does an AI agent give confident, wrong advice about your PPC?

That was my half of the session. I have been in the Amazon space for about 15 years, and I now run seven of my own brands, a 3PL, and a full service agency, so I am connecting agents to real accounts every day rather than to a demo.

My starting point is that AI made the connection easy and did not make the trust easy. Plugging Claude into Seller Central through an MCP takes minutes now. Getting advice you would actually act on takes a lot more than that.

Here is the analogy I used. A friend calls you upset: his business partner made a major decision without him and he is thinking about walking away. You tell him that is a betrayal, he should walk. Then you ask a few questions. He had been unreachable for two weeks, the decision had a Friday deadline, and his partner called him six times. Now your advice flips completely: call him back today and apologize, because he covered for you.

Same friend. Same question. Opposite advice. The only variable was context.

Your Amazon PPC agent starts every single day in exactly that position.

The agentic blind spots

Through an MCP, a PPC agent can typically see impressions, clicks, spend, orders, ACoS, CPC and conversion rate. That is it. Here is what it cannot pull from your account and will quietly assume away:

  • Customer lifetime value and subscribe and save behavior
  • Per SKU true margin after COGS, FBA fees, returns and ad spend
  • Brand pricing floors and MAP agreements
  • Inventory strategy and seasonal restock plans
  • The reason behind your target, growth mode versus profit protection versus launch
  • Seasonality, including the Q4 pull forward and the off season floor you accept
  • Ranking campaigns, where the spend is buying position rather than this week's ROAS
  • Testing phases, where a high ACoS is the plan and not the problem

Every recommendation an agent makes is only as good as what it can see. All of that list is yours to give it. If you are not sure which of your own numbers matter here, start with how to read your PPC reports.

What that costs you in practice

Take a real shape of question. "ASIN 001 is running 55% ACoS against my 25% target. What should I do?"

The answer comes back confident: cut bids 40% and pause every keyword above $1.20 CPC, which brings ACoS to target within two weeks.

The problem is that it never saw that this product sells on subscription. It optimized a single order, so it read a customer who stays with you for a year as a transaction that lost money. Cut those bids and you starve your best repeat purchase SKU.

Now ask the identical question with four extra facts attached: this ASIN has a 60% subscribe and save rate, a $42 first year customer value, and an allowable customer acquisition cost of $18.

The answer changes completely: first order ACoS is 55%, but on a lifetime basis it is 21%, well inside your $18 acquisition cost, so hold bids and put 20% more budget behind the converting terms.

The order was never the unit of profit. The customer was, and the agent could not know that until you said it. Cutting spend on a keyword that is quietly profitable is one of the most common PPC mistakes we see, and an agent will make it faster than a person will.

Write it down before you connect anything

The fix is not a better prompt. It is a context file that lives outside the chat window, in persistent memory or project instructions, so you are not retyping your business every session. Mine look roughly like this:

  • Target ACoS, with the reason: 25% because margin is 42%, not because it sounds right
  • Break even ACoS, per SKU margin map, and pricing floors including MAP
  • Which ASINs are lifetime value products and should be run aggressively
  • Current mode: launch, scale or protect
  • Max change per run: never move a bid more than 20% in a single pass
  • A never touch list: branded terms and top exact keywords, named explicitly so the agent is not guessing
  • What you will check at 7 days, defined in advance, so you can tell whether it worked

I run an agency, so I keep more than 60 of these, one per brand, plus about 500 rules across inventory, advertising and operations. Without guardrails an agent improvises on every single run. It improvises well, and that is still not how I want to run a business.

Then grade the output. Week over week, measure the results, and edit the context file as you learn. That last loop is the part most people skip. The same thinking applies to rule based bidding, which we broke down in our PPC automation playbook.

The line I would keep: the agent is the hands, your judgment is the brain.

What do 4,914 real agent conversations say sellers actually use AI for?

Akash Singh, CEO and co founder of SellerMate, did something more useful than a product demo. He read six months of real client chats with their agent, February through August 2026, and reported the pattern: 4,914 questions across 2,318 conversations from 450 clients.

The single quote that opened his deck was typed verbatim by a client: "whatever step you want to take just go ahead just optimised my campaign and implement whatever changes you want directly to my active campaign."

As Akash put it, that is not someone asking a chatbot a question. That is someone trying to hand over the job.

The twelve jobs clients hire an agent to do

Ranked by volume of questions:

Job Questions
Performance reports and audits 1,178
Keyword and search term discovery 1,022
Cutting wasted spend and negatives 675
Product and sales diagnostics 395
Bid optimization 322
Coaching and platform help 301
Dayparting and ad scheduling 269
Everything else 247
Budget management 171
Campaign creation and structure 120
Automations and rules 116
Placement analysis 98

Reporting, keywords and wasted spend dominate. Those are analyst jobs, phrased as questions rather than instructions. Automations and rules sit near the bottom.

Three stages, and almost everyone is stuck in the first two

Akash described every agent relationship moving through three stages:

  1. Answer. "What is my ACoS?" This replaces your dashboard.
  2. Diagnose. "Why did it drop, and what do I fix first?" This replaces your analyst.
  3. Operate. "Watch this every morning and fix it." This replaces the watching.

Only 23 of their 450 clients have anything running at stage three. Share of all agent interactions coming from a scheduled task went 0% in March, 0% in April, 2% in May, 11% in June, 12% in July and 21% in August. Those 23 clients now drive one in five of every interaction the agent has.

The delegation dividend

Across everything, 45% of conversations ended in a real deliverable rather than an answer in a chat window: a performance audit, a wasted spend teardown, a keyword harvest, a dayparting schedule, ready to apply.

Split that 45% average and it stops being an average. Scheduled tasks produced a finished deliverable 95% of the time. Typed ad hoc questions produced one 41% of the time. Same agent, same data. The difference is that a scheduled task already knows what it is for.

His staffing conclusions, which I think are the most quotable part of the whole session:

  • Stop hiring for watching. Nobody's job should be opening dashboards to check if something broke.
  • Start writing down judgement. The agent can execute your rules, but only the ones you have actually articulated.
  • Measure what you stopped checking. The real metric is not questions asked, it is the things you no longer look at.
  • Give it the boring 80%. Keep the calls that need context, taste, or a relationship.

One more thing Akash said that is worth hearing if you feel behind: most people are not automating end to end. The majority are still figuring out analysis. It is fine to be in stage one.

Why does the agent fail when the data underneath it is wrong?

Jon Tilley from ZonGuru took the layer below all of this. His framing: the agent does not fail, the architecture underneath it does.

His unpopular opinion, on a slide: you asked an AI agent to fix your listing, it gave you a confident wrong answer, and you blamed the agent. It was never the agent. It was what you fed it. Worse, you often do not find out it was wrong until much later than you should have.

The scale this is being built for

Jon's argument for urgency was the money moving into Amazon's AI shopping surfaces:

  • $12B in confirmed incremental sales driven by Rufus in 2025, beating Amazon's own $10B estimate
  • 350M active shoppers now using Amazon's AI shopping stack
  • 5x year over year growth in engagement with Alexa for Shopping
  • Roughly 2x growth in active users in Q2 2026 alone

His read is that Amazon has said openly it makes more money from Alexa for Shopping customers, agent payment infrastructure is live, and Accelerate landed five weeks before Q4. Where last year was a cautious test, he expects Amazon to push that surface hard through a research heavy gift buying season.

No agent can out prompt the pace

The stat that landed hardest: Google search asked marketers to absorb around 12 real shifts in 15 years. AI answer criteria are changing at roughly 342 changes per month, per an answer engine tracking firm cited by search strategist Kevin King on a panel Jon sat on.

No prompt and no hard coded rule keeps pace with that. A structured data foundation does not have to, because it is built on what stays true regardless of the rules.

Even Amazon's own AI gets confused by bad data

AVEN is Amazon's image understanding system. It does not just read the text on a listing, it recognizes what the product actually is from the photo. If your backend says V neck tee and your main image shows a crew neck, AVEN reads both, and when they contradict, your relevance score for V neck drops. You can approximate what AVEN sees for free with AWS Rekognition.

No agent failed there. The data underneath disagreed with itself, and every agent built on top of it inherited the same confusion.

The fix is engineering, not copywriting

ZonGuru's product Helix runs one pipeline in four stages: Research, Position, Structure, Perform. The agent only runs stage four. Everything that makes it right happens in the first three. On a billion dollar brand they moved an AI readiness score from 54 to 97 out of 100, and across more than 5,000 readiness reports they run, 54 is close to the average brand.

Jon also noted that a sponsored prompt inside Alexa for Shopping converts about 48% better than a typical sponsored ad, and you cannot pay your way into one. Relevance decides it.

His closing line: every agent gets better the day your data does.

I write this down every week. What we handed to an agent, what we never let one touch, and the numbers behind the call. One email most weeks, from me, and you can unsubscribe in one click. Join the Weekly Note.

What does designing listings at scale with AI agents look like?

Monte Desai from Pixii.ai covered what happens after the shopper lands: the visuals. For the broader landscape of what sellers and agencies are running, see AI tools for Amazon sellers and agencies. He has watched the full spectrum of adoption, from brands saying they will never use AI, to brands living inside Claude Code who no longer open his app at all.

The workflows he sees working:

  • A spreadsheet of ASINs handed to an agent. Each listing takes about two minutes, so a batch runs in under 20, and the output lands wherever you want it: Drive, Slack, Airtable.
  • Slack as the interface. Tag the agent with an ASIN, get seven images back in the channel.
  • Overnight queues. Build tomorrow's list, hand it over, come in and edit in the morning. Your designer does the last 20%.
  • White label inside your own tool. Agencies and SaaS platforms wiring design into their own dashboard so the client never sees the underlying app.
  • Clone a master listing across a catalogue. Perfect one listing, then scale it across variants, swapping product, ingredients, benefits, even video and A+ content. He showed 50 listings built in about 20 minutes.

On why it is worth doing at all, he cited Amazon's own research: roughly a 20% sales lift from updating a listing and keeping it fresh, because the algorithm rewards attention, and 8 to 20 percentage points of revenue from adding A+ content if you do not have it.

The point he made that I would underline for anyone worried about their team: he has not seen brands spending less or firing people. He has seen them doing far more. If a listing that cost $5,000 with an agency now costs a fraction of that, the winning move is not to cut the budget, it is to out produce your competitors.

Akash added the version of this I liked most. The output has not gone 10x, it has gone 3x to 5x, because the same person now analyzes ads on Amazon, Meta and Google instead of one platform. And the designer's job is not going away. Cookie cutter design is. Knowing the difference between good and bad output, and having the patience to iterate an agent toward the good one, is now the scarce skill.

What should a title and item highlights automation actually do?

Ben Mathew, co founder of Superfuel, closed on the most immediately actionable topic of the session, because Amazon has already changed the rules underneath every listing you own.

Amazon has split the title into two fields: a 75 character title and 125 character item highlights, rolling live through the rest of 2026. Amazon has said that titles which are not updated may be corrected by Amazon or suppressed from search results. Sellers who have executed the new titles are reporting drops in sessions, click through rate and conversion in seller forums, so this is not a cosmetic change. We wrote up the mechanics separately in our breakdown of the two part title update.

Ben's list came out of 12,546 title and item highlights edits executed and measured on live listings. If you are building this automation yourself on Claude or ChatGPT, it should:

  1. Write both fields and grade them separately. The title drives most of your impressions and clicks. Item highlights add impressions and reach conversion. Different metrics, different jobs.
  2. Name the metric it is chasing before it writes. Read where the ASIN is already strong and where it is weak, then pick the strategy that fits.
  3. Record where every keyword went. One rewrite makes six or seven decisions at once. Only a per keyword record tells you which of them paid.
  4. Rank keywords on your own numbers and competitor rank, not only search volume. Volume alone picks the right keyword 55% of the time. Your own organic rank picks it 71% of the time. Feed it your search query performance data alongside sales and traffic.
  5. Check whether your brand name earns its place. The idea that Amazon requires the brand name up front is true in a few specific categories and not a forced rule in about 90% of cases. If your brand brings fewer purchases than a category keyword, it is taking sales from your keywords.
  6. Leave some listings alone. 52% of their runs ended in no change. A suggestion for every listing is output, not a read of your catalogue.
  7. Confirm the live listing back, field by field. Amazon reports changes as executed that never appear. About 1 in 7 of theirs did not fully land. Push again, then check again, then escalate to Seller Support.
  8. Wait about a month, and drop the bad weeks. A deal or a competitor stockout moves sales more than any rewrite. The automation should account for those and then estimate the net impact of the title change itself.

Most of that also cascades to bullets, description and backend keywords. If you want the fuller picture on how the pieces fit together, start with listing optimization.

What did all five speakers agree on?

None of us coordinated, and we all landed in the same place from different directions.

  • The model is not the bottleneck. I called it the context gap. Jon called it the architecture underneath. Ben called it recording what changed. Same problem.
  • Judgement has to be written down to be executed. An agent can run your rules. It cannot run the ones that only exist in your head.
  • Scheduled beats conversational. 95% versus 41% is not a small gap. Recurring work that already knows its purpose finishes.
  • Verify what the agent claims it did. About 1 in 7 title pushes never appeared on the live listing, so a report saying the change was executed is not proof that it went through.
  • The scarce skill moved. Producing the work is no longer the hard part. Knowing which output is actually good, and having the patience to iterate an agent toward it, is what is now in short supply.

Frequently asked questions

Why do AI agents give confident but wrong Amazon PPC advice?

Because an agent connected through an MCP can typically only see impressions, clicks, spend, orders, ACoS, CPC and conversion rate. It cannot see customer lifetime value, per SKU margin, pricing floors, inventory strategy or why your target is set where it is. Without that context it optimizes a single order and gives generic advice with total confidence.

What should go in a context file for an Amazon PPC agent?

Target ACoS with the reason behind it, break even ACoS, a per SKU margin map, pricing floors including MAP, which ASINs are lifetime value products, your current mode of launch, scale or protect, a maximum change per run, an explicit never touch list of branded and top exact keywords, and what you will check at seven days.

How many sellers have actually automated Amazon work with AI agents?

Very few. Across 450 SellerMate clients and 4,914 agent questions over six months, only 23 clients had anything running on a schedule. Those 23 accounted for about one in five of all agent interactions, so adoption is real but heavily concentrated.

Are scheduled AI tasks better than asking an agent questions?

By a wide margin in this data. Scheduled tasks produced a finished deliverable 95% of the time, while typed ad hoc questions did so 41% of the time. Same agent and same data, so the difference is that a scheduled task already knows what it is for.

What is Amazon's new title and item highlights split?

Amazon has divided the product title into a 75 character title plus 125 character item highlights, rolling out through the rest of 2026. Amazon has indicated that titles which are not updated may be corrected or suppressed from search results, so it is not an optional change.

Does an AI agent actually apply the listing changes it reports?

Not always. Across 12,546 executed title and item highlights edits, about 1 in 7 did not fully land on the live listing even though the change was reported as executed. Any automation should read the live listing back field by field rather than trusting the confirmation.

What should you do this week?

If you take one thing from a 72 minute panel, take this sequence:

  1. Write one context file for one brand. Target ACoS with the reason behind it, break even, per SKU margin, pricing floors, current mode, and a never touch list. One page.
  2. Add four guardrails. Max change per run, branded terms named explicitly, what you check at 7 days, and what the agent is not allowed to do without you.
  3. Move one recurring job onto a schedule. A Monday wasted spend check or a daily account health briefing. Pick something you currently open a dashboard for.
  4. Audit your titles against the 75 and 125 character split. Then verify on the live listing that the change actually landed.
  5. Grade it after a month. Then edit the context file based on what you learned. Most people set this up once and never come back to it, and that is where the value is lost.

Assume the agent's IQ is high. The gap is context, and all of it is yours to give.