stroke 2.02px at 30pxSOUTHERLY
stroke 1.75px at 26pxSOUTHERLY
Free instrumentAlways-on Governanceabout twenty minutesin your own Claude or ChatGPT

Which of your marketing jobs should run overnight, which should wake you up, and which should never leave the room?

Not every job suits an agent that runs while you sleep. Some do: success is measurable, mistakes are cheap to undo, judgement is low. Some need a written trigger that hands them to a person before anything ships. Some should stay a conversation, because the judgement is high or nobody can put a number on what good looks like. Sorting your jobs into those three modes is the governance decision most teams skip, and it decides how much risk you carry. This diagnostic sorts your use cases, then scores how safely the always-on ones could run: tools, inbound content, data leakage and approval gates. About twenty minutes, in your own Claude or ChatGPT.

You leave with a mode table for every use case, a risk score across five vectors, a flag on any job running in a riskier mode than it should, and the first guardrails to put in.

If the link did not load the diagnostic, copy this into a new chat.

The prompt
You are a marketing operations risk assessor. I lead marketing at an Australian organisation. Help me decide which of my marketing jobs suit always-on agents, which need escalation to a person, and which should stay a conversation. Then assess how safely my operation could run the always-on ones. Be candid and practical, not alarmist.

Work in this order, one question at a time, waiting for each answer.

1. List the jobs where you use, or plan to use, AI agents. For each, tell me three things: how you would know it succeeded (a clear measure, a rough one, or judgement only); what happens if it gets it wrong (easily undone, costly to undo, or public and hard to undo); and how much judgement it needs (low, medium, high).

2. Which tools and third-party services can your agents use, and who vetted them?

3. What inbound content do your agents read (emails, reviews, comments, scraped pages, submitted briefs), and could any of it carry instructions?

4. What sensitive material could an agent reach (customer data, embargoed launches, financials), and what stops it publishing that?

5. Which agent actions happen today without a person approving them?

Sort every job from question 1 into one of three modes and explain each placement:
- Always-on: success is measurable, mistakes are cheap to undo, judgement is low. The agent runs continuously inside set limits.
- Escalated: the agent runs, but a defined trigger (a threshold, low confidence, high stakes, an exception) hands it to a person before anything ships.
- In conversation: judgement is high or success is hard to measure. The agent assists, a person decides, and it stays a dialogue.
Flag any job currently running in a riskier mode than it should.

Score five vectors from 0 to 4, where 0 is exposed and 4 is controlled:
- Mode fit: 0 jobs run always-on regardless of stakes; 2 rough sorting, triggers unwritten; 4 every job sits in the right mode with its trigger written down.
- Tools: 0 any tool freely; 2 a rough approved list; 4 a vetted set, new tools checked before use.
- Inbound content: 0 acted on unfiltered; 2 some awareness; 4 treated as untrusted, cannot trigger a high-stakes action on its own.
- Data leakage: 0 no boundary; 2 informal rules; 4 clear rules and checks on what can reach which audience.
- Approval gates: 0 agents publish freely; 2 gates on some actions; 4 gates matched to stakes, high-stakes always needs a person.

Give a total out of 20 and a one-line read of whether the operation is ready to run its always-on jobs with less supervision.

Produce the mode table, the scorecard, then the two or three guardrails to add first. Frame each as: Diagnose (the biggest exposure), Brief (the rule and the trigger to write), Make (the control to put in), Steer (who approves, what escalates, how you monitor), Compound (how the guardrail lets you safely move more jobs toward always-on).

If any vector scores 1 or below, add exactly this line: "On the ACAM and Kantar Australian AI in Marketing benchmark, AI slop and brand damage were the two most common concerns marketing leaders raised, so tightening this reduces the risk your peers worry about most."
How it runs

The diagnostic asks these in order, one at a time, and waits for each answer before scoring.

1

List the jobs where you use, or plan to use, AI agents. For each, tell me three things: how you would know it succeeded (a clear measure, a rough one, or judgement only); what happens if it gets it wrong (easily undone, costly to undo, or public and hard to undo); and how much judgement it needs (low, medium, high).

2

Which tools and third-party services can your agents use, and who vetted them?

3

What inbound content do your agents read (emails, reviews, comments, scraped pages, submitted briefs), and could any of it carry instructions?

4

What sensitive material could an agent reach (customer data, embargoed launches, financials), and what stops it publishing that?

5

Which agent actions happen today without a person approving them?

What it scores

Each dimension is scored from 0 to 4 against these anchors.

Mode fit

0 jobs run always-on regardless of stakes; 2 rough sorting, triggers unwritten; 4 every job sits in the right mode with its trigger written down.

Tools

0 any tool freely; 2 a rough approved list; 4 a vetted set, new tools checked before use.

Inbound content

0 acted on unfiltered; 2 some awareness; 4 treated as untrusted, cannot trigger a high-stakes action on its own.

Data leakage

0 no boundary; 2 informal rules; 4 clear rules and checks on what can reach which audience.

Approval gates

0 agents publish freely; 2 gates on some actions; 4 gates matched to stakes, high-stakes always needs a person.

out of 20

The three modes
Always-on

success is measurable, mistakes are cheap to undo, judgement is low. The agent runs continuously inside set limits.

Escalated

the agent runs, but a defined trigger (a threshold, low confidence, high stakes, an exception) hands it to a person before anything ships.

In conversation

judgement is high or success is hard to measure. The agent assists, a person decides, and it stays a dialogue.

The other readings
Compound Marketing

Does every piece of marketing work make the next one easier, or does each team start from scratch?

Run the diagnostic
Agent-native Operation

Can everyone on your team get the same quality from your agents, or only the person who wrote the prompt?

Run the diagnostic
Barometer

The Barometer reads your site the way an AI agent does.

Take a reading
Start hereSee the engine run.Book A Demo