TaherSoft
Limytd AG · Chief Technology Officer

Speax.ai — AI dubbing platform

Led engineering for an AI dubbing platform: agile rollout, a Go starter kit, an S3 file service, multi-threaded pipelines for translation, voice extraction and cloning, and hiring across FE/BE/AI/DevOps.

Duration
14 mo
Pipelines
3
Products on one starter kit
2
01Context

Context

Limytd AG runs two products. Speax.ai is an AI dubbing service: a customer uploads a video and gets it back speaking another language in the original speaker’s voice. Under the hood that is a chain of AI steps — transcription [draft], translation, voice extraction, voice cloning, synthesis and remux [draft] — each slow, each GPU-bound, each able to fail on its own. ChainMind is a crypto and blockchain product I had been running part-time as CTO since November 2023; in April 2024 the founder [draft] asked me to take Speax.ai full-time as well. Mandate: turn a working AI demo into a service that could take paying customers, and hire the team that would run it. Both engagements ended in May 2025.

Situation. What existed in April 2024 [draft]: AI scripts that worked when run by hand, a product idea, and no contract between the AI services and anything else. A job could run for twenty minutes [draft] and fail silently at the last step; nobody could say which step, and the only fix was to run everything again — including the GPU-expensive steps that had already succeeded. There was no project tracking, and the team still had to be hired.

The constraints that mattered: GPU time is money and minutes; the team was tiny and remote; the founder needed demos that did not break in front of customers. Success for the founder: a customer uploads and gets a result without a human babysitting the job. For the AI team: shipping a model change without touching the product.

02What I did

What I did

  1. Started from the Go starter kit. I brought the DDD + hexagonal starter kit refined since Hoitek and Armo Group: domain packages, ports and adapters, one transport layer for HTTP and the message broker, conventions written down. The backend was reviewable in week one and every hire onboarded against the conventions, not against me. Alternative rejected: Python for the backend so the AI team could own it — faster for them, but product logic would have lived in notebooks forever.
  2. Wrote the AI ↔ backend contract first. Every AI step is a worker that consumes a queue, does one thing, and publishes a result or a typed failure; jobs are idempotent by job id + step id. Broker: RabbitMQ [draft]. The AI team could redeploy a model without the product noticing; the product could add a step without touching a model. Alternative rejected: synchronous HTTP between services — simpler on day one, impossible once a step takes ten minutes.
  3. Designed the multi-threaded pipeline. A job is a state machine persisted in PostgreSQL [draft]; long media is split into segments, fanned out to worker pools with bounded concurrency per step, and joined again. Each step writes a checkpoint. In Go this is worker pools, errgroup and context cancellation [draft] — nothing exotic; the value is in the checkpoints.
  4. Built the retry patterns, manual and automatic. Manual: an operator panel with “retry from step” — replay from the last good checkpoint without re-running the GPU steps that already succeeded. Automatic: per-step attempts with exponential backoff, dead-letter queue after the budget. Manual shipped first, automatic second (the retro explains why the order matters).
  5. Wrote Bytebase. A small Go service in front of AWS S3 that owns every file moving between services: uploads, signed URLs, step outputs, lifecycle, cleanup. No service talks to S3 directly. Alternative rejected: a shared network volume — works until two clouds are involved (compute on Azure [draft], storage on S3).
  6. Hired the team. I designed the interview loop and interviewed senior frontend, backend, AI and DevOps engineers; the take-home mirrored one real pipeline step instead of a puzzle. [7] people in [3] months [draft].
  7. Introduced Monday.com and one weekly deep-dive. I taught the team the agile setup on Monday.com and ran a weekly problem-solving meeting with a single question on the table each time. It stuck for the length of the engagement [draft].
  8. Led the backend team day to day, and kept ChainMind alive part-time on the same starter kit and hiring pattern.

Results (confidence in parentheses)

  • Dubbing jobs ran unattended end to end in production (exact by construction)
  • Failed jobs needing a full re-run: from [most] to [none] after checkpoints + retry-from-step (approx, draft)
  • GPU minutes per failed job cut by roughly [40 %] because successful steps were never repeated (verify, draft)
  • [7] senior hires across FE/BE/AI/DevOps in [3] months (verify, draft)
  • Zero cross-service “file not found” incidents after Bytebase shipped (verify, draft)
  • One starter kit and one set of conventions running two products (exact, CV)
Duration
14 mo
Pipelines
3
Products on one starter kit
2
03What went right

What went right

  • The starter kit paid for itself again. Conventions written before the hires meant reviews argued about the domain, not about folder names. Credit to the two earlier teams who lived through its first versions.
  • Retry-from-step was the feature the operations side loved. “The run failed” became a one-click action instead of a chat thread [draft].
  • Bytebase removed a whole class of bugs. Once no service could touch S3 directly, “the file was there a minute ago” stopped being a sentence anyone said.
  • The take-home predicted the job. People who did well on the step exercise were good at the work; hires were productive inside their first sprint [draft].
  • The AI team got its independence. Model swaps happened without a product release.
04What went wrong

What went wrong

Blameless retro4 parts

Timeline

[draft — hypothesis] [Month 2]: the voice-cloning service went down for [an hour]. Every in-flight job failed at that step and — because automatic retry was now on — retried at once, all on the same backoff schedule. The queue filled with retries, new customer jobs sat behind them, and the cloning service came back to a wall of requests and fell over again. Noticed through a customer message about “stuck” jobs, not an alert. Understood in [a day], fixed in [three days].

Causes

  • Retry policy was per job with no global view; nothing said “a thousand jobs are retrying the same step”.
  • No circuit breaker around any AI step; the pipeline treated a dead service like a flaky one.
  • The AI service returned one error for “model crashed”, “GPU out of memory” and “bad input”, so the backend could not tell retryable from terminal.
  • Queue depth and oldest-job age were on no dashboard; alerting existed only on HTTP errors.
  • Automatic retry shipped as a feature, not as a load source; nobody modelled its worst case.
  • The weekly deep-dive cadence meant a config change waited for the meeting; there was no path for “change this today”.

Impact

[N] customer jobs delayed by [hours]; about a day of the backend team; the founder answered customer emails by hand; the AI team lost some trust in “the backend handles retries”.

Detection & response: customer message → founder → backend on-call [draft]. First response: pause the retry consumer and drain the dead-letter queue by hand. Then jitter and a retry budget per step; then the error taxonomy.

Actions

  • Jitter + a global retry budget per step — backend
  • A circuit breaker per AI step that opens on consecutive failures and pauses that queue — backend
  • retryable flag + error code in the AI ↔ backend contract — backend with the AI lead
  • Queue depth and oldest-job age on the main dashboard with paging thresholds — DevOps
  • A “pause a step” runbook anyone can run without a meeting

Second retro item [draft]: two CTO hats in parallel meant the part-time product — ChainMind — got the leftover attention. Its backend stayed at starter-kit level longer than planned and its hiring lagged. Contributing causes: no explicit split of my hours; one shared weekly meeting slot; success for ChainMind was never defined as sharply as “customers dub unattended”. Action item that should have existed: a written split of days per product, reviewed monthly with the founder.

Third retro item [draft]: compute on Azure and storage on S3 meant cross-cloud egress cost and latency nobody budgeted. Cause: GPU availability picked the compute cloud after S3 was already chosen. In hindsight, decide one cloud when the first GPU quota is requested.

05If I went back

If I went back

  • KEEPThe starter kit, the queue-per-step contract, Bytebase and manual retry-from-step, exactly as they were
  • CHANGEShip the error taxonomy and the circuit breaker before switching on automatic retries — and automatic only with a budget
  • CHANGEQueue depth and oldest-job age on a screen on day one, before the first customer
  • CHANGEOne cloud, decided together with the GPU quota
  • CHANGENo second CTO hat without a written split of days
06Lessons

Lessons

  • A pipeline is only as debuggable as its checkpoints — the checkpoint is the product, the step is the detail.
  • The interview exercise should be a slice of the job, not a puzzle.
  • “Part-time CTO” of a second product is a full-time cost to the first one unless it is written down.

More with Go

Speax.ai — AI dubbing platform · TaherSoft