Stop reading the model's work out loud. The brief is the management moment. The progress feed is not. Stronger models perform better when you define the finish line, the acceptance test, the constraints, and the permission boundary up front, then review the finished artifact against that standard. OpenAI's Codex prompting guide tells builders to remove prompts for upfront plans, preambles, and status updates during rollout because those instructions can make the model stop before the job is finished (OpenAI). Anthropic's Opus 5 guide makes the same point from the other side: it performs best when given the full task specification up front and left to run (Anthropic).
The Habit That Now Gets Worse Results
A marketing director watches an agent draft a launch note, interrupts to correct phrasing, asks for progress updates, and steers the work paragraph by paragraph.
Anthropic's usage research puts numbers on that split. In real Claude Code sessions, humans make about 70% of the planning decisions and only about 20% of the execution decisions, while a single user prompt triggers around 10 model actions on average and sometimes more than 100 (Anthropic). The human contribution is deciding what to build, which tradeoffs matter, and what counts as done.
OpenAI's current model guidance lines up with that pattern. It tells builders to describe the expected outcome, success criteria, allowed side effects, evidence rules, and output shape, then avoid step by step process guidance unless the exact path matters (OpenAI). GPT-5.6 guidance adds the permission layer: define what level of action the request authorizes so the model can continue in scope work without unnecessary pauses while still stopping before destructive or scope-expanding actions (OpenAI).
The Five Fields That Belong in the Brief
If you want a model to do bounded work without live steering, put these five fields in the initial brief and nothing ornamental around them:
- Finish line: the concrete artifact that must exist at the end.
- Success criteria: the standards that decide whether the artifact is acceptable.
- Constraints: the rules the model cannot violate.
- Permissions: the actions it is allowed to take without asking again.
- Blocking questions only: the narrow conditions that justify interrupting the run.
Here is the structure in practice:
- Finish line: Deliver one approval-ready landing-page draft and one approval-ready paid-social variant for the July campaign.
- Success criteria: Use Magnet voice, include three sourced claims from the approved research set, stay between 700 and 900 words for the landing page, write one CTA, and surface one unresolved risk if a claim cannot be supported.
- Constraints: No unsupported figures, no invented client claims, no publication, no outreach, no legal conclusions, and no sources outside the approved URL list.
- Permissions: You may read the approved sources, draft locally, rewrite for clarity, and spend up to $25 in tool or API cost. Escalate before any external posting, purchase, destructive edit, or scope expansion beyond the campaign deliverables.
- Blocking questions only: Interrupt only if the audience, offer, legal claim, or approved source set has more than one plausible reading that would produce materially different work.
Narration Is an Audit Artifact, Not a Live Dashboard
This does not mean reasoning traces have no value. OpenAI's monitorability research found that, across 13 evaluations and 24 environments, monitoring chains of thought was substantially more effective than monitoring actions and final outputs alone (OpenAI). That supports logging and automated monitoring.
It is not a case for live human management by reading every thought as it appears. Anthropic's faithfulness research found that Claude 3.7 Sonnet mentioned a decision-changing hint only 25% of the time, while DeepSeek R1 did so 39% of the time (Anthropic). The narration is not a reliable live map of what drove the result. Use it for automated checks, incident review, and post-hoc audit.
Anthropic's Opus 5 guide makes that operational. The model narrates readily during agentic work, so teams are told to tune its update cadence and keep updates brief and meaningful (Anthropic). OpenAI goes further and says to remove prompts for mid-run status updates entirely in its Codex harness guidance because they can break completion (OpenAI).
Inspect the Artifact and the Gate
Anthropic's cloud version of Claude Cowork is designed around unattended execution with a final review and approval gate before anything ships (NBC News; TechCrunch). Google Ads' July 2026 terms coverage makes the same point in a regulated marketing context: the advertiser retains the obligation to review, approve, or remove automatically generated campaigns and ad assets (PPC Land).
That is the durable pattern for marketing teams. Write the brief like a contract. Let the model run inside that contract. Review the delivered artifact against the contract. Approve, revise, or reject.
There is also a cost reason to stop watching the feed. Microsoft Research's SentinelBench found that polling agents cost 5.1 times more at 10 minutes and 9.7 times more at 40 minutes than event-waiting agents, while completing fewer tasks (Microsoft Research). Put a spend bound in permissions and escalate only when the brief says the run crossed a real threshold.
The Management Standard
Most teams need better briefs. BCG reports that 96% of CMOs say they are pursuing significant end to end AI transformation, yet only about 8% are connecting multiple agents to run campaigns autonomously, and 42% still use generative AI only as an individual-task assistant (BCG). The gap is operating discipline.
If you want stronger model performance, stop managing the scroll. Manage the brief. Define the finish line. Write the success criteria. Lock the constraints. State the permissions. Limit interruptions to blocking questions only. Put the spend or escalation bound inside permissions and inspect the artifact at the gate.
That is what management looks like now.
Work With Magnet
Magnet helps marketing teams turn AI use from a loose prompt habit into a governed production system with bounded briefs, approval gates, and review logic that holds up under real workloads. If your team is still supervising narration instead of grading artifacts, talk to Magnet.
Sources
- OpenAI model guidance
- OpenAI GPT-5.6 guidance
- OpenAI Codex prompting guide
- OpenAI chain-of-thought monitorability research
- Anthropic Opus 5 prompting guide
- Anthropic Claude Code expertise research
- Anthropic reasoning-faithfulness research
- NBC News on Claude Cowork cloud runs
- TechCrunch on Claude Cowork cloud runs
- PPC Land on Google Ads July 2026 terms
- Microsoft Research SentinelBench
- BCG on agentic marketing transformation


