BUILD VS BUY · APR 20, 2026 · 8 MIN

How to Evaluate an AI Agent Partner for Your Marketing Stack: A CMO Checklist

Most marketing leaders have been pitched by an AI agent development company in the last six months. Most of those pitches are not the right fit for a social commerce stack. Here is the checklist we use with CMOs evaluating agent partners.


The phrase "AI agent" is doing a lot of work in 2026. It is being applied to everything from a single LLM call inside a Zapier workflow to a multi-agent autonomous system that reads, reasons, and writes back to your CRM. When a marketing leader gets pitched by an "AI agent development company," the first job is to figure out which of these actually applies to the work, and which is product-marketing language that obscures the build.

We have run vendor evaluations for CMOs at brands ranging from $20M to $400M in revenue. The patterns are consistent. This piece is the checklist we use, written for a marketing leader who is comparing two or three agent partners and trying to make a defensible decision.

What an "AI Agent" Should Mean in a Marketing Context

For a marketing org, a useful AI agent is a piece of software that does three things. It takes an input (a SKU launch, an inbound creator request, a paid-media performance report). It runs a sequence of LLM calls and tool executions against that input. It produces an output that a human reviews and acts on (a draft brief, a draft listing, a recommended creative variant).

Note what is not in that definition. There is no "fully autonomous." There is no "replaces a team." There is no "self-improving." Those are pitch words. The agents that work in production are bounded, reviewable, and produce drafts that humans approve.

If a vendor pitches an agent that does not have a clear human-in-the-loop checkpoint, that is a flag. Either the agent is doing trivial work (in which case you are paying agent prices for automation), or it is doing important work without review (in which case it will fail in a way that is expensive).

The Six Checklist Items

Use these six questions on every vendor pitch.

1. Where Does the Agent Run, and Who Owns the Code?

Three configurations we see in the market:

For a marketing org, the third configuration is almost always the right answer. The prompt library and the data integrations are too core to outsource, and the build cost difference is small relative to the multi-year operating cost.

2. What Data Does the Agent Need, and Where Does It Live?

A useful agent needs access to your PIM, your customer data, your creative library, and your performance reporting. Most marketing data lives in a fragmented stack: Shopify, a PIM (or no PIM), GA4, Meta Ads, TikTok Ads, an affiliate platform, a CRM.

The agent partner you want is the one that asks about your data architecture in the first conversation. The agent partner you do not want is the one who has a "we connect to everything" answer that turns out to mean "we have a Zapier integration that breaks on the first edge case."

For most marketing orgs, the right pattern is: clean the PIM and the customer data first (covered in the marketing data foundation piece), then layer the agent on top. Vendors who do not understand the data layer are not the right partners.

3. What Is the Build Time, and What Is the Operate Cost?

Honest ranges for a marketing-org agent build:

A vendor pitching a six-week build for $400K is overpriced. A vendor pitching a one-week build for $20K is underbuilding. The middle of these ranges is where the working agents in our network sit.

4. What Does the Hand-Off Look Like?

The hand-off is when the build is done and the agent moves into production. This is the moment where most agent programs fail. The vendor finishes, the team does not know how to operate the agent, the prompt library degrades, and within 12 months the agent is broken.

Ask the vendor: what does the hand-off package include? The right answer covers documentation, prompt library version control, monitoring dashboards, escalation playbook, and a 90-day post-launch support window. The wrong answer is "we can do a lunch-and-learn."

5. How Do They Handle the Brand Voice Problem?

The agent has to produce output that sounds like your brand. This is where most agent vendors fail in marketing contexts. Their prompt library was trained on a generic enterprise voice, and the output reads like every other brand on their roster.

Ask to see five examples of output the vendor has produced for current clients. If the outputs sound interchangeable, the prompt library is not where it needs to be. The good vendors have a brand voice intake process: they study your existing copy, they do a voice workshop with your content team, they iterate the prompt library against samples before launch.

6. What Are the Reference Calls?

Always do three reference calls. Ask each reference: what was the build experience, what did the agent actually produce in the first 90 days, and what is the operate cost looking like at the 12-month mark. The answers will tell you more than any pitch deck.

The reference call to watch out for is the one where the customer is enthusiastic about the build experience but vague about the production outcomes. That pattern usually means the agent looks great in demo and underperforms in operation.

Red Flags

Five patterns that have been associated with bad outcomes.

The vendor cannot show the prompt library. "It is proprietary" is the wrong answer. You are buying the prompt library. You should see it.

The pitch leans on autonomy. Phrases like "self-improving," "fully autonomous," or "replaces your team" are red flags. The working agents are bounded and reviewable.

No mention of the data layer. If the vendor has not asked about your PIM, your CRM, or your performance data within the first hour, they are selling a feature, not a system.

Pricing is opaque. Working partners give you a clear range. Pitch decks with "investment varies based on scope" and no numbers are usually expensive.

No post-launch operate model. The build is the easy part. The 12 to 36 months of operating the agent is where the value compounds or evaporates.

Build vs Buy in 2026

For marketing-org agents specifically, the right answer in most cases is "buy the build, own the operation." A vendor builds the initial agent, hands over the code and prompt library, and your team operates it. The build is a one-time cost. The operate model is what determines long-term ROI.

The brands building agents fully in-house (with no vendor) are typically large enough to have an internal AI engineering team. The brands buying fully managed agent services are typically not yet sure what they want and are using the vendor as the figure-it-out partner, which is expensive but sometimes correct.

Most CMOs we work with land in the middle: vendor-built, customer-owned, vendor-supported for the first year. The total cost over three years is meaningfully lower than fully managed, and the internal capability that builds up is the long-term asset.

We cover the broader generative AI consulting evaluation framework in a separate piece. For agent-specific evaluations, the checklist above is what we use.

What This Looks Like When It Works

A DTC apparel brand we worked with ran this evaluation in late 2025. They had three vendors in final round. Two pitched fully managed agent services at $300K to $500K per year. One pitched a build-then-operate engagement at $140K for the build plus $60K per year for support. The third was the right answer.

Twelve months in, the brand has a listing generation agent producing 1,200 listings per month, a brief generation agent producing 600 briefs per month, and a creative variant agent producing 200 variants per week. Total cost over twelve months including vendor support and internal operator time was $310K. GMV produced from agent outputs was roughly 14x that.

The framing for the CMO conversation is not "should we buy an AI agent." The framing is "we are going to operate AI agents in our marketing stack for the next decade. Which partner sets us up to operate them well."

If you want help running this evaluation against current vendor pitches, the diagnostic at clankersapp.com is where to start.