What an AI agent can and cannot automate in a content team
The honest split between the work an agent should take off your content team and the work that has to stay with a person.
Drafted by the same agent we build for clients. Reviewed and approved by a human.An agent can take the repeatable half of a content team's week: gathering source material, first drafts against a brief, formatting, filing, checking a new piece against what you already published. It cannot take judgement, accountability, or a voice nobody has written down yet. The split is not creative versus mechanical.
The quick version:
- An agent can hold any step with a checkable right answer.
- Judgement, accountability and undocumented voice stay with a person.
- The human approval gate belongs in the design from the first sketch.
- A good first workflow is frequent, written down, and low risk.
Where the line actually sits
The line sits at one question: does the step have a checkable right answer? If a reviewer could look at the output and say yes or no against something already written down, an agent can do the first pass. If the honest answer is that it depends who we are this quarter, it stays with a person.
Creative versus mechanical is the wrong axis. Plenty of creative work is mechanical, and plenty of admin needs judgement. The content system we built for a marketing and design startup is drawn along that one question, and it is live in production.
| Hand it to the agent | Keep it with a person |
|---|---|
| Pulling source material together before a draft exists | Deciding what is worth saying this quarter |
| First drafts written against a written brief | Judging whether the draft sounds like you |
| Formatting, tagging, filing, cross linking | Sign off, and the accountability that comes with it |
| Checking a new piece against everything already published | Deciding which brand risk is worth taking |
Nothing in the right column is work your team wanted to give up. The left column is what they do at six in the evening because it has to get done, and it is the half that quietly sets how long everything takes.
Two questions that sort any step in a minute
Two questions get you most of the way. First: could someone write the rule for this step down in an afternoon? If the rule only exists as taste, an agent has nothing to check itself against.
Second: if this step goes wrong, who finds out first, us or a customer? Steps where the company catches its own mistakes are the safe ones to hand over.
The approval gate is where the design starts
The approval gate is the first thing to design, and everything else gets arranged around it. It is tempting to leave approval as a checkbox bolted on once the build works. That order is backwards, because the gate decides how much of the work you can safely hand over at all.
How the gate works in a live Slack system
The gate sits inside a system running in Slack for a marketing and design startup of about 60 people. Four agents split the work. One runs the thread, one retrieves source material and is the only one allowed to read Notion, one writes, and one reviews the draft against the sources and returns approve or revise.
That review loop runs between the agents, so nobody reads a raw first draft. Then the run stops. The agent posts in the thread and asks, and nothing continues until a person answers.
When the answer comes back, the agent files the deliverable and a task log to Notion, a step by step record of what it did, and marks the card ready to publish. Publishing stays a human action. The system is live in production.
Why a gate you can switch off is not a gate
A gate that can be switched off in a busy week is a suggestion. The one in this system is the run itself: the agent cannot proceed without an answer, so there is no fast path that skips review when a deadline is close.
Take the gate out and the only thing that changes is when you find out something went wrong. It moves to after the piece has shipped, with your name on it.
Three kinds of work to keep off the agent's list
Three kinds of work stay with people even once the gate is in place. Each fails the checkable-answer test for a different reason, and it helps to know which reason applies before you argue about it in a planning meeting.
The reply when something has gone wrong
The reply after a mistake, a late delivery or a public complaint has one job: to show that a person took responsibility. An agent drafting it removes the one ingredient that made it worth reading. A human writes it, and the agent can pull the context they need.
Claims you would have to defend
Claims you would have to defend belong to whoever can stand behind them. Pricing, positioning, results, anything comparative about a competitor. Those need someone who knows what the company has actually promised and who will be in the room when a customer pushes back on it.
Voice that does not exist as text yet
Voice that has never been written down cannot be copied. A new brand, a new sub brand, a founder's own writing before anyone has described what makes it sound like them. An agent matches a documented voice well. It cannot infer one from taste, and asking it to teaches your team that the output is always wrong.
Describing the voice in writing is work a person does once. It is also the part that decides whether a first build is worth doing, which is why it comes up early in an Agent Audit.
Five tests for picking your first workflow
Five tests tell you whether a workflow is a sensible first build. Size matters less than shape, so ignore how small the candidate looks and run it through all five.
- It runs weekly or more often.
- The steps are the same every time, and someone could write them down in an afternoon.
- The output already passes a reviewer today, so the gate fits a step that exists.
- A mistake surfaces internally before it reaches a customer.
- Nobody on the team enjoys doing it.
The second test fails in two different ways, and they point in opposite directions. If the steps genuinely change every run, this is not your first build.
If the steps are steady but nobody has written them down, write them down first. The build gets much easier afterwards, and the write up is worth having even if you never commission one.
If a candidate passes and the question becomes whether to build or to hire, the cost comparison sits in its own post.
FAQ
These are the questions content leads ask before they commit to a first build.
Can an agent write our blog posts end to end?
An agent can draft a blog post end to end, but end to end stops at the approval gate. It works against a brief, checks itself against the sources, and hands you something to approve or send back. Publishing stays a human click.
What happens when the agent gets something wrong?
When the agent gets something wrong, a person sees it before anyone outside the company does. The gate holds the run until someone answers, and the task log in Notion shows every step it took. You then correct the brief or the sources rather than guessing at the model.
Will this disrupt my team for months?
Disruption is the thing this setup is built to avoid. The client side of the work is a couple of Looms walking through the process as it happens, one live session of about 45 minutes, and the rest async. Nobody has to learn a new tool, because the agent lives in the tools they already use.
How do I know which workflow to automate first?
You find the workflow to automate first by running the five tests above over one week of real work. The candidate that passes all five is the one to write down. If nothing passes, the tests your candidates failed tell you what to fix first.
What does a first engagement with Arrox cost?
A first engagement with Arrox starts at the Agent Audit, priced at €500 for a 90 minute call and a two page memo, credited forward against later work within 30 days. You get the money back if the memo is not useful. The rest of the ladder sits on the offer page.
Do we need to change tools to run this?
Changing tools is not part of it. The system runs in Slack and Notion because that is where the team already works, and an agent parked in a tool nobody opens is easy to ignore. The live case study shows the shape of it.
Running the five tests at your next content standup
Running the five tests costs you one standup. Read them out over the work your team is doing that day, and mark which ones each candidate passes.
If something passes all five, write its steps down in plain language: the trigger, the inputs, each step, who checks it, what done looks like. That document is what a builder works from, and it is most of the input to a first build.
If nothing passes, the test each candidate failed tells you the next move. A candidate that fails on frequency points you at a different workflow. One that fails because mistakes reach customers before anyone sees them needs a review step first, which is a process fix rather than a build.
Either way you leave the standup with the next move written down: a candidate to document, or the specific test to fix first. If you want a second pair of eyes on the shortlist before you commit budget, that is what the Agent Audit is for.