From Ghost to EmDash: moving my blog onto Cloudflare Workers

This blog ran on Ghost for five years. It now runs on EmDash, on Cloudflare Workers, with the words in D1 and the images in R2. Every old URL still resolves. Here is how the migration actually went — the pipeline, the four bugs that fought back, and what I would do differently.
Why move at all
Ghost did its job. It served the blog, it had a decent editor, and it never lost a post. What I wanted was fewer moving parts: no MySQL instance to keep patched, no server quietly accumulating CVEs, and a publishing path I could reason about end to end. I also wanted to modernise the reading experience, which had grown a bit tired.
The replacement had to be cheap to run, boring to operate, and — the constraint that shaped everything else — it had to keep every existing URL working.
What had to survive the move
- 32 posts, one standalone page, 66 images and 12 tags.
- Every indexed URL, including the trailing-slash form the old site used.
- The RSS feed, the sitemap, and the ads and analytics already in the pages.
- Original publication dates, so the archive still reads in the right order.
URLs are a promise. A migration that quietly breaks your links is not a migration, it is a rewrite with extra steps.
The stack I landed on
The site is Astro with output: "server" — every page is rendered on request, so publishing does not require a build. EmDash provides the content model and the admin. Underneath it is a single Cloudflare Worker with a few bindings:
- D1 for content, schema and settings.
- R2 for media.
- KV for the object cache and sessions.
- The Images binding for resizing, plus a cron trigger for scheduled work.
The interesting property is what is missing: there is no database server, no origin host, and no build pipeline between me hitting publish and the post going live.
Getting the content out of Ghost
The obvious route is Ghost's Content API, which is documented and returns JSON. On this install it returned 404 for every version prefix I tried, including the ones that should work on 4.x. I could have spent the day on it. Instead I scraped the public HTML, which turned out to carry everything I needed.
The archive pages give slug, title, date, reading time, excerpt, tags and feature image. Each post page gives the body, and an exact article:published_time in the metadata. That is enough to rebuild the content faithfully.
// Two sources, joined on the slug.
const cards = await readAllArchiveCards(); // /page/1, /page/2, ...
const posts = [];
for (const card of cards) {
const page = await readPostPage(card.slug); // body + published_time
posts.push({ ...card, ...page });
}Turning five years of HTML into Portable Text
EmDash stores rich text as Portable Text — a JSON array of typed blocks rather than a blob of HTML. So the import needed a converter, and the mapping is mostly mechanical:
Source HTML | Portable Text |
|---|---|
h2–h6 | block with a style |
p | block, style normal |
ul / ol | block with listItem and level |
pre > code.language-x | code block with a Prism language id |
table | table block with rows and cells |
figure > img | image block with a media reference |
iframe embeds | htmlBlock |
Lists needed the most care: Ghost nests them inside the parent list item, so nested lists have to be lifted out into their own blocks with an incremented level. Callout cards became blockquotes, and embeds are kept as raw HTML blocks so nothing is silently dropped.
The whole corpus came to 498 text blocks, 44 code blocks, 9 tables and 40 images. Downloads were de-duplicated by content hash, so shared images were stored once.
The four bugs that ate the most time
The import ran on the first try and produced content that looked right. It was not right. These four took the rest of the day.
1. Images that rendered without a file extension
A Portable Text image block needs its asset to be a full media reference — type, ref, url and provider — not just the media id. With only the ref set, the renderer cannot resolve the file, falls back to using the id itself, and produces a URL with no extension that 404s.
// Broken: the renderer cannot resolve this
"asset": { "_ref": "01M3SE55R9JX7D7V3DJFKSYV80" }
// Correct: a media reference
"asset": {
"_type": "reference",
"_ref": "01M3SE55R9JX7D7V3DJFKSYV80",
"url": "/_emdash/api/media/file/01M3SE55HYK7GHRG6GXQNHN8VX.png",
"provider": "local"
}39 of 64 images were broken across 13 posts. The clue was that featured images were fine — image fields accept a bare id, so only the inline body images were affected.
The more useful lesson is about verification. My first check said there were no broken images. It was wrong, because the regular expression I used to find image URLs required a file extension — which was precisely the thing that was missing. A test that assumes the shape of the thing it is checking is not a test.
2. Dark mode that followed the OS and ignored the button
The theme switcher puts a class on the html element and sets color-scheme; the colour tokens are written with light-dark(). Clicking the button set the class, set color-scheme, and changed absolutely nothing.
The CSS minifier was lowering light-dark() into a pair of custom properties that flip inside a prefers-color-scheme media query. That mechanism cannot be influenced by a class, so the palette was pinned to whatever the operating system said. Declaring modern browser targets fixed it — which is the sort of thing that works perfectly in development and only misbehaves once the production pipeline gets hold of it.
3. A grid that was exactly 48 pixels too narrow
The article layout is three columns: a meta rail, a centred body column, and a sidebar. At a 1440px window the body measured 712px instead of the 760px I had asked for.
max-width is measured from the border box, so a max-width that precisely sums the grid tracks still has to account for the element's own padding. The design was right; the arithmetic was off by two lots of 24px.
4. A badge that was never fetched
One post embeds a vendor button as an image. It was in the DOM, correctly sized in the markup, and invisible. The image carried loading="lazy", and a lazy image whose layout box is zero-height by zero-width never gets fetched — the browser has nothing to measure against the viewport. Removing the attribute was the whole fix.
Deployment held a surprise too
The Cloudflare adapter snapshots wrangler.jsonc into dist/server/wrangler.json at build time, and the deploy command ships that generated file rather than the source you just edited. I added a variable to the source and redeployed without rebuilding, so the site shipped with an empty configuration block and the setup wizard refused to run. If you change Worker configuration: rebuild, then deploy.
Keeping every URL
The old site served every page with a trailing slash, and those URLs are indexed. Astro can enforce that with trailingSlash: "always", except that it also 404s the CMS's own routes — including the endpoint the admin needs to generate types. So the setting is "ignore", and the legacy shape is preserved with a 301 in middleware instead.
The middleware skips the admin, the asset paths, anything that looks like a file — and the .well-known directory. That last one matters: OAuth and MCP clients fetch discovery documents at the exact path, and a redirect breaks the lookup. I found that one the hard way, by breaking it.
The result is 57 URLs — every post, tag, archive page and feed — checked automatically and returning 200.
Publishing did not behave the way I expected
Creating an entry with a publication date stores it as a draft; publishing always stamps the current time. For a migration that means two steps: create the entry with its content and date, then publish it, passing the original timestamp so the archive keeps its chronology. A direct column update follows as a backstop, because getting this wrong rewrites five years of history.
What it costs to run
One Worker, one database, one bucket. For a blog this size that sits comfortably inside the free tiers, and the only thing I actually patch is a package version. The trade is that I now own the theme — which is the point, and also the cost.
What I would do differently
- Reach for the source before fighting the API. The scrape took twenty minutes; the API archaeology had already taken longer.
- Test against a deployed site much earlier. Two of the four bugs above only exist after the production CSS pipeline and the edge cache get involved.
- Write the verification first, and make it check what actually renders — fetch every page, extract every image URL, request each one.
- Use a long-lived API token from the start. The CLI credentials expire after an hour, and a migration takes longer than an hour.
Where it ended up
The blog is faster, the writing workflow is unchanged, and every link people have saved over the last five years still goes where it always went. The content import took an afternoon. The last ten percent — the pixels, the caches, the configuration — took considerably longer, which is the normal shape of these things.



