Most teams do not decide to scale a web application. They find out they need to, usually when a dashboard that loaded in 200 milliseconds last quarter now takes four seconds, or a marketing campaign sends a spike of signups that the database cannot absorb. The instinct in that moment is often to talk about a full rewrite on a "modern" stack. In practice, almost every application can absorb an order of magnitude more traffic through a sequence of targeted changes, long before a rewrite is justified.
This guide walks through how to tell when scaling work is actually needed, which fixes usually come first, and where the real tradeoffs sit between caching, database scaling, horizontal scaling, and architectural changes like splitting a monolith into services. It is written for technical founders and product leaders who need to make this call with their engineering team, not for engineers who already live in this problem daily.
What "scaling" actually means
Scaling a web application means keeping response times and reliability stable as the number of users, requests, or data volume grows. It is not one technique. It is a set of independent levers, and most production incidents trace back to one specific lever, not the architecture as a whole:
- Request volume: more concurrent users hitting the same endpoints.
- Data volume: tables and indexes growing past what fits comfortably in memory or gets scanned efficiently.
- Compute-heavy operations: report generation, file processing, or AI inference that blocks normal request handling.
- Write contention: many processes trying to update the same rows at the same time.
Before changing anything, find out which of these is actually the bottleneck. Teams that jump straight to "we need microservices" or "we need to move to Kubernetes" without this step often spend months on infrastructure changes that do not touch the real problem, which is usually one slow query or one synchronous call to a third-party API sitting in the critical path.
Start with measurement, not architecture
Application performance monitoring, database slow-query logs, and basic infrastructure metrics (CPU, memory, connection counts) will usually show the actual constraint within a day of looking. Common findings, roughly in order of frequency:
- A handful of database queries without proper indexes, each fine individually but expensive at scale.
- N+1 query patterns, where a list view triggers one query per row instead of one query for the whole page.
- Synchronous calls to external APIs or payment processors sitting inside the main request path.
- No caching layer at all, so every request recomputes something that barely changes.
- A single application server or database instance sized for the traffic from a year ago.
Fixing the first two items alone resolves a large share of "scaling" problems without touching infrastructure at all. This matters because infrastructure changes carry real cost and risk, and they are not reversible in the same afternoon a bad index fix is.
Vertical scaling: the fastest short-term lever
Vertical scaling means giving an existing server more CPU, memory, or faster storage instead of adding more servers. It is the simplest option because the application code and deployment process do not change, and it buys time to make better decisions under less pressure.
Its limits are real, though. There is a ceiling on how large a single machine can get, it creates a single point of failure, and cost increases faster than capacity once you are on larger instance sizes. Vertical scaling is the right first move when traffic is growing steadily and predictably. It is the wrong move when traffic is spiky, because you end up paying for peak capacity around the clock.
Horizontal scaling: more servers behind a load balancer
Horizontal scaling means running multiple copies of the application behind a load balancer, so traffic spreads across several smaller servers instead of one large one. Cloud providers' guidance on performance efficiency treats this as the default pattern for handling variable demand without overpaying for idle capacity (see the AWS Well-Architected Framework's performance efficiency pillar).
Horizontal scaling only works cleanly if the application is stateless: any server should be able to handle any request without relying on data stored only in that server's local memory or disk. Common blockers to statelessness include:
- Session data stored in server memory instead of a shared store like Redis.
- File uploads written to local disk instead of object storage (such as S3 or an equivalent).
- In-process caches that get out of sync between server instances.
- Background jobs tied to a specific server rather than a shared job queue.
Fixing statelessness is usually a smaller project than it sounds, and it is worth doing before adding more servers. Once an application is stateless, autoscaling (adding and removing servers automatically based on load) becomes straightforward with most cloud platforms.
Caching: the highest return-on-effort fix
A caching layer stores the result of an expensive operation so the next request for the same thing is nearly instant. It has the best ratio of effort to impact of any item on this list, because most applications recompute the same results repeatedly for different users.
Three layers are worth considering, usually in this order:
- HTTP and CDN caching for static assets and public pages, so requests never reach the application server at all. The caching model behind this is described in the HTTP caching documentation on MDN.
- Application-level caching with a store like Redis, for query results, computed values, or rendered page fragments that change infrequently. Redis's own operational guidance describes how to scale a cache horizontally once a single instance stops being enough (see Redis scaling documentation).
- Database query caching, either at the application layer or through a database-native mechanism, for the specific slow queries measurement already identified.
The hard part of caching is not the technology, it is invalidation: deciding when a cached value is stale and needs to be refreshed. Underinvesting in invalidation logic is how teams end up with user-facing bugs where someone sees data that is minutes or hours out of date after an update. Plan cache expiry and invalidation rules at the same time as the caching layer itself, not as an afterthought.
Scaling the database
The database is usually the hardest part of a system to scale, because unlike application servers, it holds state that has to stay consistent. Three approaches are common, and they solve different problems:
- Read replicas are read-only copies of the primary database that handle reporting queries, dashboards, and other read-heavy traffic, freeing the primary instance to handle writes. PostgreSQL's own documentation on streaming replication covers how this mechanism works at the database level (see PostgreSQL's warm standby and streaming replication documentation). Read replicas help with request volume but do not reduce the total amount of data a single instance has to store.
- Connection pooling reduces the overhead of opening and closing database connections for every request, which becomes a real bottleneck once an application runs dozens of parallel server processes.
- Sharding splits data across multiple database instances by some key, such as customer ID or region, so no single instance holds the full dataset. Sharding solves data volume problems that replicas cannot, but it adds real complexity: cross-shard queries and reports become harder, and a poorly chosen shard key can concentrate traffic on one shard anyway.
Sharding is usually the last resort, not the first move. Most applications never need it. Query optimization, proper indexing, read replicas, and archiving old data out of hot tables solve the overwhelming majority of database scaling problems.
Asynchronous processing: taking work out of the request path
Anything that does not need to finish before a response is sent back to the user, such as sending a confirmation email, generating a PDF, resizing an image, or calling a slow third-party API, belongs in a background job queue rather than the main request path. This single change often fixes perceived "scaling" problems that are actually just slow synchronous work blocking fast requests behind it.
A job queue also absorbs traffic spikes gracefully: incoming work queues up and processes steadily instead of overwhelming the application or database all at once.
When a monolith actually needs to become services
Splitting a monolithic application into separate services should solve a specific, demonstrated problem, such as one component needing a different scaling profile, a different release cadence, or a different technology than the rest of the system. It should not be done because microservices are the current trend.
The real cost of splitting a system is operational: more deployments to coordinate, more network calls that can fail, and the need for proper service-to-service observability so a slow or failing service is visible before it takes down everything that depends on it. For most applications under meaningful but not extreme load, a well-optimized monolith with caching, a scaled database, and horizontal scaling behind a load balancer will outperform a half-finished microservices migration, both in reliability and in cost.
A practical order of operations
When traffic or data growth starts causing real problems, this order tends to produce the best result for the least risk:
- Measure first. Identify the actual bottleneck with APM tools and slow-query logs.
- Fix the obvious: missing indexes, N+1 queries, synchronous third-party calls.
- Add caching at the layer where it will have the most impact, with a clear invalidation strategy.
- Make the application stateless if it is not already.
- Scale the database with read replicas and connection pooling before considering sharding.
- Scale the application horizontally behind a load balancer, with autoscaling if traffic is spiky.
- Only split into services when a specific component has a scaling or operational need the rest of the system does not.
Teams that follow roughly this order usually find that the "we need a rewrite" conversation never has to happen. The underlying architecture from the original web application stays largely intact, and the work is incremental rather than a multi-month replatforming project with its own risk of introducing new bugs.
How this plays out differently for a SaaS product
Multi-tenant SaaS applications add one more dimension: scaling has to account for noisy-neighbor effects, where one customer's heavy usage should not degrade performance for everyone else. This usually means rate limiting per tenant, careful indexing on tenant ID, and sometimes isolating the largest customers onto dedicated infrastructure. Teams building or scaling a SaaS platform should plan for this from early on, since retrofitting tenant isolation after customers are already on a shared, unthrottled system is considerably harder than designing for it upfront. The backend framework you started with also shapes how far a monolith can stretch before any of this becomes urgent, which is worth weighing when choosing a backend framework for a SaaS product.
Frequently asked questions
How do I know if my application actually needs to scale, or if something is just broken?
Check whether performance degrades specifically under higher load (more users, larger datasets, peak hours) or whether it is slow consistently regardless of traffic. Consistent slowness, even with low traffic, usually points to a specific bug or missing index rather than a scaling problem, and should be fixed before any infrastructure changes.
Is a database migration to a "bigger" database engine a valid scaling strategy?
Rarely on its own. Most relational databases (PostgreSQL, MySQL, and their managed cloud equivalents) can handle far more load than most applications generate, once queries are optimized and indexed correctly. Switching engines without fixing query patterns usually just delays the same problem.
Should a small application plan for scale before it has traffic?
Plan for statelessness and reasonable database design from day one, since both are cheap to get right early and expensive to retrofit. Hold off on load balancers, caching layers, or service splits until there is real traffic data showing they are needed. Premature infrastructure for traffic that never arrives is a common source of wasted budget.
How much does it cost to scale a web application?
It depends heavily on which levers are actually needed. Query and index fixes are mostly engineering time with little added infrastructure cost. Caching layers and read replicas add modest, predictable monthly infrastructure cost. A move to microservices or a major database re-architecture is the most expensive path, both in infrastructure and in engineering time, which is exactly why it should be the last option considered, not the first.
Getting the sequence right
The teams that handle growth well are not the ones with the most advanced infrastructure. They are the ones that measured before building, fixed the cheap things first, and only reached for the expensive architectural changes once the data showed they were necessary. If your application is starting to struggle under load and you want a second opinion on where the actual bottleneck is before committing to a rewrite or a major infrastructure project, book a strategy call or get in touch to talk through what is actually happening in your system.