Hoivalani — home-visit healthcare
CTO for a home-visit and family healthcare product: event storming, a team of 4 devs, 3 designers and 2 nurses, DDD + hexagonal Go microservices on RabbitMQ with CQRS, and a React nurse panel with live location tracking.
- Team size
- 9
- Duration
- Approximate date15 mo
- Hand-written API clients
- 0
Context
Hoitek is a Finnish company building Hoivalani, a product for home-visit and family healthcare: nurses drive to patients’ homes, and the company needs to schedule those visits, know where each nurse is, and keep the records straight. I joined as CTO in early 2022 [verify], remote from Tehran, reporting to the founder [verify], with a mandate to take the product from [idea / prototype — verify] to a working platform and to build the team that would run it.
Situation. When I arrived there was [no backend / a prototype — verify], a design team that needed direction, and a domain none of the engineers had lived in. The biggest constraint was domain knowledge: healthcare scheduling has rules nobody writes down. The second was that we were a small remote team across [Finland and Iran — verify] and needed a process that didn’t depend on being in one room.
What I did
- Put the nurses in the room, literally. I ran event-storming sessions on Miro with two nurses as team members, not stakeholders we interviewed once. We mapped the visit lifecycle — request, plan, drive, arrive, treat, report, bill [verify] — before writing a line of code. Alternative rejected: writing user stories from the founder’s description; it would have been faster and wrong.
- Chose Go with DDD and hexagonal architecture. The domain was rule-heavy and would change as we learned; hexagonal kept the rules away from transport and storage. Services talked over RabbitMQ with a transport layer that let us swap sync/async without touching the domain. Alternative rejected: a Node monolith — faster to start, but I’d seen it become unmaintainable at [Lemon / earlier job — verify] once the domain got rich.
- CQRS + Redis for the read side. Schedules and live nurse state are read hundreds of times for every write, so queries went through Redis-backed read models. Cost: eventual consistency confused the panel for a while (see the retro).
- Wrote the OpenAPI → React Query codegen. Backend code produced OpenAPI specs; a tool turned them into typed React Query hooks. The frontend never hand-wrote an API call. This is the piece I’ve carried to every project since.
- Docker for every environment, self-hosted GitLab CI/CD and registry. Dev, test and prod were the same images. Alternative rejected: GitLab.com SaaS — [data residency / cost — verify].
- Built the nurse-management panel myself (React, RTK, MUI, i18next): staff and shift cycles, live nurse location over WebSocket, charts, and a Leaflet map for dispatch.
- Directed the designers through all Figma pages so design stayed one step ahead of engineering.
Results (confidence in parentheses)
- Platform in production with [N] nurses and [N] patients in [city/region] (verify)
- Team hired and onboarded: 4 developers, 3 designers, 2 nurses (exact, CV)
- [N] services on RabbitMQ, [N] releases per week through CI (verify)
- Zero hand-written API clients on the frontend after the codegen shipped (exact by construction)
- Live location and state for every active visit in the panel (exact, CV)
- [Uptime / incident count / any number the founder would sign — verify]
- Team size
- 9
- Duration
- Approximate date15 mo
- Hand-written API clients
- 0
What went right
- Domain immersion paid for itself. Having nurses as teammates killed at least [three — verify, draft] features before they were built. Same lesson as cooking fifty dishes at CookThePot.
- The codegen changed the team’s rhythm. Backend merged a spec, the frontend got types in minutes, and integration bugs [dropped to near zero — verify, draft].
- Hexagonal held up. When [the scheduling rules changed — verify, draft], the change was in the domain package only.
What went wrong
Timeline
[draft — hypothesis: keep, correct or replace] [Month 3 — verify]: the panel started showing stale nurse states after writes. Noticed by a nurse during a demo, not by monitoring. Root cause understood in [2 days — verify], fix in [a week — verify].
Causes
- CQRS was introduced before the team had a shared model of eventual consistency; the panel assumed read-after-write.
- Read-model invalidation lived in three places with no single owner.
- We had no end-to-end test that crossed the write side and the read side; the CI pipeline tested each service alone.
- Remote, cross-timezone reviews meant the invalidation PRs were merged by whoever was awake.
Impact
[One — verify] demo went badly; the nurses lost some trust in the “live” label; roughly [a sprint — verify] went to the fix and to re-explaining the model.
Actions
- Read-model invalidation moved into one module with one owner✓ Done
- An integration test suite ran the RabbitMQ path in CI✓ Done
- The panel showed “updating…” states instead of pretending to be synchronous✓ Done
- CQRS was documented in an ADR the whole team reviewed✓ Done
Second retro item [candidate — verify or delete]: the OpenAPI tooling was built before the API stabilised, so early spec changes forced tooling changes; in hindsight, build the tool once two services have stable contracts.
If I went back
- KEEPStart with event storming
- KEEPChoose Go + hexagonal architecture
- CHANGEIntroduce CQRS one aggregate at a time, with the team watching the first one, instead of as a day-one architecture decision
- CHANGESet up the end-to-end test path before the second service exists