All work

Platform & tooling

CI Translation Components

Turning a hackathon prototype into a production translation pipeline that ships localized content as ordinary merge requests.

Role
Author and overall DRI
Organisation
GitLab
Period
2026
One CI component, grounded in specifications, a termbase and translation instructions, turning changes in Orbit, about.gitlab.com and GitLab Operator into Japanese, Korean and French merge requests, each reviewed by GitLab Duo

What

Translation at most companies lives outside the tools engineers actually use. Content is exported, mailed to a vendor, translated somewhere else, and pasted back in weeks later — by which point the source has moved on. The result is localized content that is permanently, structurally stale.

CI Translation Components started as my submission to GitLab's 2026 AI Enterprise Hackathon: a set of reusable GitLab CI/CD components that treat translation as a pipeline stage rather than a side quest. A source file changes, CI notices, the content is translated, and the translation arrives as an ordinary merge request in the same repository — reviewable, revertible, and auditable like any other change.

The trial was deliberately narrow. The goal for v1 was not to solve every localization use case; it was to prove that a CI component could reliably handle a small, well-bounded set of real production content and create the right merge requests without an engineer standing over it. Building something customers could adopt meant using it ourselves first.

I wrote the roadmap and carried the epic as overall DRI, sharing ownership with the Linguistic Quality Assurance DRI who owned translation quality.

How

Pipeline shape. Source change → detect → translate → commit → merge request → auto-merge. Every model call routes through the GitLab AI Gateway; the runner never talks to a vendor API directly, which is what made the whole thing defensible from a security review perspective.

Quality starts before the model sees a word. Nothing is translated cold. Every job is grounded in three inputs: a set of specifications for what to translate and how, a termbase maintained in GitLab, and translation instructions that set tone, style and locale conventions. Because all three live in version control, terminology and voice are reviewable, diffable and owned — just like the content they shape.

Review stays adjacent, not blocking. This was the design decision the trial turned on. If a linguist has to sign off before a translation MR can merge, throughput collapses to the speed of human review and you have rebuilt the vendor bottleneck inside CI. Instead, the first review pass happens where the translation lands: GitLab Duo reviews the merge request as soon as it arrives, and on its own it does a genuinely solid job. Linguistic review then fires after merge — a Translation Review Flagger flow opens async tracking issues in a separate linguistic-review-tracker project — so quality problems get found and fixed without holding the pipeline hostage.

The quality bar is configurable, not fixed. Teams that need a higher bar have plenty of levers, all native to GitLab: custom MR review instructions for Duo, additional CI jobs that gate on their own checks, Vale rules that enforce style and terminology, and the usual approval rules. The component sets a sensible default; each project decides how tight to hold the line.

Blind spots before production. Before touching anything customer-facing, two parallel workstreams ran the components against forks — Spanish Solutions pages on an about.gitlab.com fork, and /docs on a GitLab Operator test fork. Forks first meant the failure modes we found were free. The one hard prerequisite was unglamorous: group-level AI_API_KEY and GITLAB_API_TOKEN provisioning, which blocked every production pipeline until it landed.

First production wave. Four projects, chosen because they were real but bounded:

ProjectContentLanguages
about.gitlab.com44 Solutions pages across 7 categoriesSpanish
GitLab Operator11 documentation pages, excluding /developerKorean, French, Japanese
Orbit29 documentation pages, the full docs setJapanese
GitLab releases3 release posts, 19.2–19.4, simshipped with EnglishJapanese

Orbit was the proof of continuous: I translated its entire documentation set into Japanese and kept it current as the source changed, with translation MRs typically landing about 20 minutes after the source commit.

Linguists shipping releases, end to end. The biggest proof point wasn't a page count — it was who was driving. I set our linguists up to configure these pipelines themselves and self-serve translations from source change to merged MR, with no engineer in the loop. They used it to simship GitLab's 19.2, 19.3 and 19.4 releases in Japanese, live on the same day as the English original rather than weeks behind it. That is the shift the project was built for: localization owned by the people who own the language, running on the same platform as everything else.

Rollout in three phases. Blind-spot identification, then production enablement, then full operational coverage — each with a date and a named owner rather than a vague "when it's ready".

Security posture up front. Least-privilege access for the component and its service accounts, secrets masked and hidden, every run traceable from request through pipeline to MR, and a rollback path that any DRI could execute immediately. Writing these down before rollout is what let the trial move fast afterwards.

Outcomes

The trial ran its window and the epic closed in August 2026. What it produced:

  • CI Translation Components running in production on real projects, with translations landing as merge requests in the repositories that own the content.
  • A documented, repeatable workflow from request → CI execution → translation MR → review → merge, with named owners, monitoring and support paths — the thing that was missing when this was just a hackathon demo.
  • Runbooks aimed at customers, not just us: how to adopt the component, how to troubleshoot a failed job, and how to request or contribute changes. The point of dogfooding was always to hand it onward.
  • A clear path forward. The trial ended with a written stabilize and expand epic rather than a victory lap, which is the honest outcome: production taught us about token rotation pausing forks, post-processing that needed hardening, and configuration policy that had to be decided deliberately rather than inherited. Each lesson is captured as scoped, documented work, so whoever picks the components up next starts from a plan and a working production baseline — not a blank page.

The targets the trial was measured against, set before it started: pipeline success rate of at least 50% of production runs completing without manual engineering intervention, MR creation succeeding for at least 80% of successful runs, and a median turnaround from source change to open MR inside one business day, with minutes as the stretch goal. Orbit's Japanese docs routinely hit the stretch goal, turning around in about 20 minutes.