Skip to main content
Penyaskito Blog

Main navigation

  • Home
Language switcher
  • English
  • Español
User account menu
  • Log in

Breadcrumb

  1. Home

Can it be all so simple? Migrating a 13-year-old WordPress blog to Drupal CMS

By penyaskito , 10 October, 2026
Image
Can It Be All So Simple logo

Can it be all so simple? Spoiler: no, it can't. But it was fun.

Can It Be All So Simple (CIBASS for friends, and yes, the name comes from the Wu-Tang Clan song) is a Spanish blog about cinema, music, comics, TV series and culture that has been publishing since 2013. Some friends started it, and later a bunch of others joined, including me. I ended up hosting it and doing its maintenance for years. And we even ended up earning an award.

It ran on WordPress all that time. Maintenance was painful. In 2021, I alleviated part of that pain by migrating it to roots/bedrock and managing updates with Composer, but still, I wanted to migrate it to Drupal at some point. If you explore its contents, you will end up finding quite a few articles written by Drupal people. 

On October 1st, 2026, thanks to Drupal CMS 2.2.0 and its multilingual support improvements, that could finally happen. Even if activity on the blog has been at its lowest, I still wanted to do that to keep the site online with minimal maintenance, while I refuse to accept that the project is almost dead. 

What we were migrating

The first step: an inventory. By setting up both the WP project and the new Drupal CMS one in DDEV, we made sure they could see each other. Looking at what was actually in the database:

  • 717 published posts (and 5 drafts), 2013 to 2022.
  • 4,110 attachments. 3,741 of them JPEGs.
  • 545 approved comments, 104 pingbacks, and 1,190 spam comments that stayed behind.
  • 13 categories and… 5,861 tags. For 717 posts. Yes.
  • 3 users. Bylines lived in the post body as "Por NAME, @HANDLE".

Audit your content before the migration too, not only after. The database will surprise you.

In 2021 I said that upgrading your site is the best time for taking the trash out. Still true: the spam and 5,580 revisions stayed in WordPress.

The shape of it: Drupal CMS, recipes and a direct database source

The site is built on the Drupal CMS project template. Everything specific to CIBASS lives in project recipes: one per vocabulary, one per content type, one per media type, all bundled in a site template recipe. A fresh clone installs the whole site, in Spanish, with a single command. Then the migrations fill it.

For the source I didn't use the XML export of wordpress_migrate. wordpress_migrate_sql reads the WordPress database directly. With DDEV, the WordPress clone runs as its own project, and the Drupal one connects to its database container. 24 migrations, run in order by a script that stops at the first failure.

The development happened on a parallel branch of the same repository, and at cutover it was a plain git merge into master, plus running the migration live. The WordPress code is now only in git history. RIP.

The content model: let the data decide

WordPress had "posts". CIBASS had three different things pretending to be posts: critiques, interviews and articles. So the plan was to split them, and the first plan said "around 337 critiques".

The data said otherwise. A critique on CIBASS is a post with a rating, and the rating is… an image. A picture of a score at the end of the post. So the rule ended up being "a critique is a post with exactly one rating image". Posts with several ratings are compilations ("our 10 favorite movies of the year"), and those became articles flagged for editorial review.

The final numbers: 185 critiques, 22 interviews (plus 2 manual overrides) and 513 articles. Ratings are now a real field, and a custom field formatter outputs that same image.

Bylines became a contributor vocabulary. Parsing "Por NAME, @HANDLE" sounds easy until you find multiple authors, authors joined with " y ", and the same person written in 61 different ways. A small overrides table at migration time brought it down to 46.

Twelve rules to read a blog post

This is where most of the work went. Thirteen years of hand-written HTML, three generations of WordPress editors (classic, shortcodes, Gutenberg), and a lot of copy-paste from YouTube. While at it, we wanted to use the power of media in Drupal, so we needed to ensure we could unify all these different kinds of embeds from several different social networks this site has outlived (hey Vine!).

As an example, every critique body goes through a chain of twelve process plugins, in this order:

field_content:
  - plugin: wp_strip_byline_blocks
  - plugin: wp_rewrite_inline_formatting
  - plugin: wp_strip_rating_image
  - plugin: wp_rewrite_gallery_shortcodes
  - plugin: wp_rewrite_manual_galleries
  - plugin: wp_rewrite_caption_shortcodes
  - plugin: wp_rewrite_gutenberg_blocks
  - plugin: wp_rewrite_embed_shortcodes
  - plugin: cibass_image_override
  - plugin: wp_rewrite_img_urls
  - plugin: wp_relativize_links
  - plugin: wp_rewrite_search_to_tag

Each one does one thing. Remove the byline, because it's a field now. Remove the rating image, same reason. Turn [gallery ids="…"] (62 of them, 828 images) and the galleries people built by hand into gallery media. Turn 521 [caption] shortcodes into embedded media with captions. Turn [embed], bare URLs and raw iframes into the right remote media. Turn every <img src="…/uploads/…"> into a <drupal-media> token. Make links relative. And convert old ?s=query search links into tag pages.

Order matters. If the rating image isn't stripped (step 3) before images are converted (step 10), the rating image ends up as a media item in the body.

And then there's the rule that is not a rule: CIBASS editors use [sic], [mos] and [habla] as editorial brackets in Spanish. They look exactly like shortcodes. They are not. Leave them alone.

Write each transformation as its own small process plugin, with its own tests. You'll rerun the migration dozens of times, and you want to know which rule broke.

That's around 2,250 lines of process plugins and 186 test methods in the migration module. It sounds like a lot. It's the reason I could rerun everything without fear.

Media: eleven types for one blog

Those rules need somewhere to put things. The site ended up with eleven media types, and some of them deserve a story:

  • Images (4,102): caption and description moved from the post to the media item, so they travel with the image.
  • Remote video (344, YouTube and Vimeo): 73 of them were dead. Some got curated replacements; the rest went into an editorial backlog, created automatically during the migration, for human review.
  • Gallery (65): a custom media source that holds a list of images. A gallery is a media type referencing other media types 🤯.
  • Animated image (24): GIFs get their own type that skips image styles. Core's AVIF image effect re-encodes through GD, which only keeps the first frame. Your animated GIF stops being animated, silently.
  • Remote audio (Spotify, SoundCloud, iVoox): iVoox doesn't have a usable oEmbed endpoint, so a custom resource fetcher provides one. Old embed.spotify.com iframes needed to be parsed and rewritten too.
  • Remote document (Scribd, SlideShare): Scribd's oEmbed endpoint returns a 406 unless you explicitly ask for format=json.
  • Social (X, TikTok, Instagram): Meta's oEmbed requires business verification. So Instagram embeds use the blockquote markup and their embed.js instead.

oEmbed Providers made the custom providers manageable, as core's list doesn't include iVoox or Scribd. And every third-party embed goes through a formatter that waits for Klaro consent before loading the iframe. No request reaches Spotify or YouTube until the visitor says yes. So in the process, we made the recipe GDPR-friendly.

Media types are prerequisites, not follow-ups. If a post is migrated before its media type exists, its embeds are gone and you'll run the migration again.

I planned remote video as "something for later". It wasn't.

Don't break the internet

Thirteen years of URLs are out there: in Google, in tweets, in other people's posts. As I mentioned, the site even won an award, and it's still getting lots of organic traffic. Breaking URLs is not an option.

  • Post URLs keep the WordPress pattern (/2015/01/23/las-claves-de-akira), built by Pathauto from the local post date. Use the GMT one and some posts move to the previous day.
  • Old slugs: 54 redirects for posts whose slug changed over the years.
  • Attachment pages: WordPress creates a page for every image, and those were around 4,100 URLs, most of what Google had indexed. Instead of 4,100 redirect entities, one event subscriber sends them to the parent post. I verified it against 1,966 real URLs.
  • Uploads: old /app/uploads/… paths, checked against 4,109 real files and their resized variants. 284 of them had non-ASCII filenames, because of course they did.
  • Old sitemaps: nginx sends the WordPress sitemap URLs to /sitemap.xml. For RSS, WordPress provided several different paths for the same feed! We redirect all of them to the proper one.

Verify redirects with real requests, not with migration counts.

Owning our data

Here's a fun one. The sidebar had a "most visited posts" widget, powered by Jetpack Stats. Thirteen years of visits… that never left WordPress.com. We had no way to export them. When WordPress goes, the stats go with it. The new block started with a list copied from a Wayback Machine snapshot of the live site. Archaeology instead of analytics.

That was a good reminder of something we keep saying in open source and keep forgetting in practice: if it's not on your servers, it's not yours.

CIBASS now uses Matomo, self-hosted on our own server. It runs cookieless, and doesn't send any data to third-party companies. And unlike with Jetpack, this time we could take our history with us: our Google Analytics data since 2023 is now imported into Matomo, so we didn't start from zero.

As of this week, that "most visited" block ranks posts by real pageviews: cron asks Matomo's Reporting API once a day for the most viewed pages in the last 3 months, maps them to posts by path alias, and stores the ranking. Without depending on any external provider.

Now the parts we care about most — the content, the code, the deployments and the numbers — live on infrastructure we control. The code lives on a self-hosted GitLab, too.

When you migrate away from a platform, check what data you can't take with you. Then make sure it doesn't happen again.

Numbers

  • WordPress repo: 84 commits, 2021–2022 (mostly maintenance). Drupal: 337 commits, April–October 2026.
  • 722 posts (drafts included) → 185 critiques, 24 interviews, 513 articles.
  • 545 comments and 104 pingbacks migrated. 1,190 spam comments didn't make it.
  • 24 migrations, 12 body-rewriting rules, 11 media types.
  • About 4,100 legacy URLs handled by one rule.

What's next

One thing we lost in the move was our mailing list, provided by Jetpack. I'm often frustrated with organizations or content creators building a community or audience, and having all that data depending on third-party services. If you don't have the email addresses, you don't have a community. If we revive the project, we will need to start that from scratch. Hopefully Simplenews can help with that.

Conclusions

Managing WordPress is hard if you want traceability of the updates in version control. roots/bedrock helps a lot with that. While they do a great job maintaining an endpoint for Composer, it means depending on a third-party company.

Jetpack provides great functionality, but you lose all sovereignty over your data. It ends up on WordPress.com servers, not yours.

With this move we got most of our data back, maintenance will be easier from now on, and our editors get a unified UX thanks to the media work. All on an infrastructure I know well, and that is proven on other sites I maintain.

Have you migrated a WordPress site to Drupal CMS? Any experiences worth sharing?

AI was used for a drafting an outline of this post. 

Tags

  • Drupal
  • Drupal CMS
  • migration
  • WordPress
  • Drupal planet
  • Add new comment

Comments2

The content of this field is kept private and will not be shown publicly.

Plain text

  • No HTML tags allowed.
  • Web page addresses and email addresses turn into links automatically.
  • Lines and paragraphs break automatically.

Manuel Garcia (not verified)

52 min ago

Fantastic work

"at cutover it was a plain git merge into master, plus running the migration live"

Very nice 👏

  • Reply

Sr. Arbogast (not verified)

1 sec ago

Opinión

Gran artículo. Enhorabuena.

  • Reply

Monthly archive

  • October 2026 (2)
  • July 2026 (3)
  • April 2026 (4)
  • August 2025 (1)
  • April 2025 (1)
  • July 2023 (1)
  • December 2021 (1)
  • May 2021 (2)
  • April 2021 (1)
  • September 2014 (1)
  • November 2012 (1)
  • September 2012 (2)
  • August 2012 (3)
  • June 2012 (6)

Recent content

Can it be all so simple? Migrating a 13-year-old WordPress blog to Drupal CMS
9 hours 41 minutes ago
Quarterly Contributions summary for 2026 Q3
5 days 8 hours ago
Canvas Tips&Tricks: Declaring images in SDCs
2 months 3 weeks ago

Recent comments

Opinión
3 hours 8 minutes ago
Fantastic work
4 hours ago
Created project/canvas…
2 months 2 weeks ago

Blogs I follow

  • Mateu Aguiló "e0ipso"
  • Gábor Hojtsy
  • Pedro Cambra
  • Can It Be All So Simple
  • Maria Arias de Reyna "Délawen"
  • Matt Glaman
  • Daniel Wehner
  • Jacob Rockowitz
  • Wim Leers
  • Dries Buytaert
  • arcturus
  • Drupal Core AI digest
  • Drupal CMS AI digest
  • Drupal Canvas AI digest
  • Drupal AI AI digest
  • Drupal Patterns AI digest
  • Trisha Gee
  • Très Bien Tech, by _nod
  • Moshe Weitzman
  • Drupal core change records
  • Ed Zitron's Where's Your Ed At
  • Sebastian Bergmann (phpunit.expert)
  • PHP Reads
  • mandclu
  • Marissa Epstein
  • WordPress Core
  • David Bushell
  • Tag1
Syndicate

Footer

  • Drupal.org
  • LinkedIn
  • GitHub
  • Mastodon
  • Twitter
Powered by Drupal

Free 🇵🇸