Draftly, a multi-agent AI content platform

2026 to present · SynergyBoat · published

Content teams needed blogs, carousels, short videos and social copy without a studio behind them. Built Draftly: seven agents take a blog from brief to image plan in a fixed order inside a durable workflow, each step naming the capability tier it needs while the router walks a failover chain up from there, and the run ends in a revise loop that rewrites flagged blocks until the score clears or three rounds are up. Live at draftly.synergyboat.com, with a 13-tool MCP server beside it.

Numbers this essay already carries, each linking the paragraph that states it
server tests 1,639 stated in
suite run time 26 seconds stated in

The problem

A model will write you a blog post. Ask it for twenty and you get the same post twenty times: the same hedged opening, the same three-item list in the middle, the same closing paragraph that summarises what you just read. The tool that writes the draft is the worst available judge of whether the draft sounds like a person.

So the work worth doing is holding a brand’s voice across pieces, keeping claims tied to source material rather than invented, scoring the result against something other than the model’s own opinion, and then producing whatever shape the channel takes. Draftly makes five of those shapes: a blog, a carousel, a short video, post copy for LinkedIn and X, and an X thread. Each has its own path through the system, and the blog’s is the long one.

What shipped

A multi-tenant platform, self-serve from signup: Next.js on Cloudflare Workers, Bun and Hono on EC2, Postgres with pgvector, BullMQ for queued work, Remotion for video, and a 13-tool MCP server so an agent can score, revise and save content without opening the app at all.

A blog runs as a durable workflow of eight steps, and the third of them runs seven agents in a fixed order: brief-creator, research, drafter, seo, editorial-polish, humanizer, image-planner. Research drops out when the piece already carries a source extract and its stored reference URLs are under thirty days old, so a run is seven agents or six. Then the tail. The draft is scored, and while the score still carries flagged signals the flagged blocks are rewritten and scored again, up to three rounds. After that, metadata, the render of the blocks to markdown, and the run closes.

The server suite is 1,639 tests across 236 files and it finishes in 26 seconds on my laptop, which is the number that decides whether I run it before a change or after one. The one usage figure the repository holds is 117 completed runs between April and June 2026, written into my own cost roadmap on 2 July. It came from a database query that is not in the repository, so I can quote it and I cannot re-run it, which turned out to be a finding of its own.

A tier is a capability floor. It says nothing about price.

How it works

Every step of the plan carries a number from 1 to 5, and the numbers are set by hand: brief-creator asks 4, research 3, drafter 5, seo 3, editorial-polish 4, humanizer 3, image-planner 4 (Figure 1). The router reads that number as a floor. It keeps every model at or above it whose provider key is configured, sorts by tier and then by a hand-set priority, and wraps the result in a failover chain, so a rate limit on the first model drops the call to the next one. If a step throws, the runner retries it one tier higher, capped at 5.

A tier is a capability floor. It says nothing about price. No part of the routing reads a cost table. Tier 5 holds a flash model at $0.30 per million input tokens beside a Sonnet at $3.00, and the sort cannot tell them apart. Cost did move two of the floors, but I moved them: one commit in July 2026 took the humanizer from tier 4 to 3 and editorial polish from 5 to 4, because the humanizer makes more calls than anything else in the pipeline and that is where the spend was.

The planner is not a model either. It builds the same ordered list on every run from a template in code, with the single rule that drops research, and a comment above it puts the ordering constraint in capitals: humanize is the terminal pass. A model does plan the edits a reader asks for after the first draft, picking the agents, the order and a tier for each. That is the only place in the system where planning is inference.

Nor is the score at the end a model’s verdict. Eight measured layers, weighted into one number out of 100, decide it: lexical, structural, rhetorical, epistemic, cadence, quality, substance and redundancy, with the revise threshold at 70. The humanizer’s middle layer measures burstiness, lexical diversity, Writer’s Diet and em-dash density on a block, and only spends a model call on the blocks that fall below target. A rewrite that comes back having invented a specific is rejected.

What I’d do differently

The ladder assumes more than one provider underneath it, and the deployment has one. CI writes the server’s environment from a Gemini key and an OpenAI key; the production file the repository points at has the Gemini key filled and the OpenAI one empty. In August I removed the only Gemini model sitting at tier 3, in a commit whose message says “unused”. Put those two together and the chains for tiers 3, 4 and 5 are one chain: every planned step starts on the flagship model, and an escalation from 3 to 4 re-runs the step on the model that just failed it. I cannot read the deployed secrets from the repository, so this is what the repository shows rather than what production certainly does, and that gap is the point. Nothing anywhere reports the chain a step actually resolved to. The floors read as a ladder in the code, in the plan and in the logs, and under one key they are a flat surface. The registry should refuse to start, or at least say something, when a tier a step asks for holds no model of its own.

The second thing is that cost is measured in a way nobody can audit. A ledger collects calls and tokens for the length of a run, prices them off a table the file itself calls an estimate, writes one log line and clears. No table in the database carries a token, cost, tier or model column. The two tail steps that call a flagship model, the revise loop and the metadata pass, sit outside the ledger entirely. And because the failover wrapper reports its first link’s model, a run that fell through to the second one files its tokens under the first one’s name. So the figure I quote for a blog, roughly $0.23 before three fixes and $0.15 to $0.17 after, is an estimate that prices every stage at the flagship, and I would trade all of it for one run measured end to end.

Third, the README. It described a nine-stage pipeline two weeks before the server it describes existed, and it was still describing it this month. The facts I gathered to write this essay list nineteen claims in my own README, docs and code comments that the code contradicts, and three of them had reached this page. Documentation does not rot on a day you can point at, which is why I counted the disagreements instead of trusting a read-through.

One thing about how it was built, because it explains the shape. Another engineer scaffolded the Next.js app in March; the other 1,330 of the repository’s 1,335 commits are mine, and 1,134 of all 1,335 name a Claude model as a co-author. Work arriving at that rate is exactly the condition under which a README quietly stops being true. It is the argument for the 1,639 tests, and against trusting the prose beside them.

The Draftly tier floors and the failover chain Seven planned agent steps run left to right, and each one asks for a capability tier from 1 to 5 as a floor, set by hand in the plan. The floors, in order: brief-creator 4, research 3, drafter 5, seo 3, editorial-polish 4, humanizer 3, image-planner 4. Below them, the chain the router walks upward from a floor: it keeps every model at or above the floor whose provider key is configured, sorts by tier and then by a hand-set priority, and wraps the result in a failover chain. With the openai, gemini and anthropic keys, a step at minTier 3 resolves to gpt-5.4-mini, then gemini-2.5-pro, then gemini-3.5-flash, then claude-sonnet-4-6, then gpt-5.4, and a step at minTier 4 or 5 resolves to gemini-2.5-pro, then gemini-3.5-flash, then claude-sonnet-4-6, then gpt-5.4. With the gemini key only, minTier 3, 4 and 5 all resolve to the same chain, gemini-2.5-pro, then gemini-3.5-flash. When a step throws, the runner retries it one tier higher, capped at 5. With three keys that reaches a different chain. With the gemini key alone the chains are identical, so the retry re-runs the model that just failed. No cost lane is drawn, because the repository records no cost per stage. each step asks for a tier floor 5 4 3 4 3 5 3 4 3 4 brief- creator research drafter seo editorial- polish humanizer image- planner the floors are hand-set in the plan, one per step a floor is the lowest tier the router may pick, never a price every model at or above it whose provider key is set, lowest tier first with the openai, gemini and anthropic keys minTier 3 gpt-5.4-mini gemini-2.5-pro gemini-3.5-flash claude-sonnet-4-6 gpt-5.4 a step at tier 3 starts on the mini model when an openai key is set minTier 4 and 5 gemini-2.5-pro gemini-3.5-flash claude-sonnet-4-6 gpt-5.4 tiers 4 and 5 already share one chain, even with three keys with the gemini key only minTier 3, 4 and 5 gemini-2.5-pro gemini-3.5-flash one key flattens the ladder: all three floors resolve here escalate one tier on a thrown error, capped at 5 the same chain, so the retry re-runs the model that just failed no cost lane: the repository records no cost per stage
part what it does
brief-creator Asks for tier 4 as a floor.
research Asks for tier 3 as a floor.
drafter Asks for tier 5 as a floor.
seo Asks for tier 3 as a floor.
editorial-polish Asks for tier 4 as a floor.
humanizer Asks for tier 3 as a floor.
image-planner Asks for tier 4 as a floor.
minTier 3 with the openai, gemini and anthropic keys: gpt-5.4-mini, then gemini-2.5-pro, then gemini-3.5-flash, then claude-sonnet-4-6, then gpt-5.4.
minTier 4 and 5 with the openai, gemini and anthropic keys: gemini-2.5-pro, then gemini-3.5-flash, then claude-sonnet-4-6, then gpt-5.4.
minTier 3, 4 and 5 with the gemini key only: gemini-2.5-pro, then gemini-3.5-flash.
escalate A step that throws is retried one tier higher, capped at 5. With three keys that reaches a different chain.
one key With the gemini key only the three chains are identical, so the retry re-runs the model that just failed.
no cost lane Nothing in the routing reads a price, and the repository records no cost per stage.
Fig. 1. The tier floors and the failover chain. Each step asks for a floor, and the router takes the lowest-tier model at or above it whose key is set. Two key sets are drawn: with three keys a tier-3 step starts on a mini model, and with the gemini key alone every floor resolves to the same chain, so escalating a tier re-runs the model that just failed. Drawn for Draftly only.