When PayStream first came to us, they had a classic startup problem: their monolithic Ruby on Rails application had carried them to 500,000 users and $50M in annual transaction volume, but every new feature deployment was a gamble. Database locks during peak hours, 30-second page loads on the dashboard, and a deployment process that required a 2AM maintenance window were becoming existential threats to the business.
We decided against the popular 'rewrite everything in microservices' approach. Instead, we employed the Strangler Fig pattern, gradually replacing components of the monolith while keeping the system running. The first service we extracted was the payment processing engine, moving it to a Go-based service with event sourcing. This alone reduced transaction latency by 85% and eliminated the database lock issues.
The database strategy was crucial. We migrated from a single PostgreSQL instance to a sharded architecture using Citus, keeping PostgreSQL's familiar query interface while distributing data across multiple nodes. For the read-heavy dashboard queries, we introduced a CQRS pattern with materialized views updated through a Kafka event stream. Dashboard load times dropped from 30 seconds to under 200 milliseconds.
Caching was implemented at every layer. We used Redis for session management and hot data, a CDN for static assets, and application-level caching with cache invalidation driven by domain events. The cache hit ratio reached 94%, meaning only 6% of requests actually hit the database.
The results spoke for themselves. PayStream scaled from 500K to 10M users over 18 months without a single significant outage. P99 latency dropped from 12 seconds to 180 milliseconds. Deployment frequency increased from bi-weekly to multiple times per day. The engineering team grew from 8 to 35 developers, and the new architecture made onboarding new engineers significantly faster.
Topics