Haystack

How to measure an AI marketing agent

Zaki Hasan

An agent can produce more work and make marketing worse. The evaluation needs to reward useful selection and safe completion while exposing how much human attention the system still consumes.

The decision this method improves

This method determines whether an agent deserves broader scope, continued supervision, a tighter rule, or removal from a job. Results should be reviewed by job type because a content publication and a recipient-facing email have different risks and feedback cycles.

A working method

  1. Define the intended result and quality standard for each job before activation.
  2. Measure opportunity acceptance separately from draft approval.
  3. Categorize factual, voice, structure, relevance, policy, and destination corrections.
  4. Track approval wait, execution success, retries, failures, suppressions, and manual stops.
  5. Connect observable outcomes such as live pages, useful replies, placements, citations, and qualified visits.
  6. Record reviewer time and compare it with the previous manual process.
  7. Expand autonomy only when quality and outcomes hold with less supervision.

The evidence to keep

Retain the opportunity rationale, generated asset, policy applied, reviewer decision, external action reference, timestamps, error state, and later outcomes. External analytics and CRM systems should remain sources of truth for traffic, product behavior, and revenue.

The common failure mode

The easiest metrics to collect are drafts, pages, sends, and impressions. Optimizing those can reward weak opportunity selection and high-volume behavior. Another failure is assigning perfect causal credit to a page or message in a multi-touch buying journey.

How to measure the result

Use a scorecard with acceptance, approval, correction, cycle time, execution reliability, safety interventions, qualified outcome, and supervision cost. Compare within the same job and risk class over time.

What to do next

Choose one job and write its quality rubric and outcome definition before reviewing last month's actions. Haystack's attribution guide explains how observable events attach to shipped work.

Frequently asked questions

What is the best single metric for an AI marketing agent?
There is no sound single metric. Use completed useful work with a quality threshold, then report downstream contribution and human supervision cost beside it.
Should rejected work count as failure?
Track rejection reasons. Early rejections can be valuable if they improve the rule and prevent poor external actions. Persistent repeated rejection indicates the job is not ready for autonomy.
How is agent attribution different from revenue attribution?
Agent attribution connects work with nearby observable outcomes. Revenue attribution also needs customer and pipeline data and should state uncertainty across a multi-touch journey.