Service Mesh vs API Gateway for Microservices Traffic
API gateways handle external traffic; service meshes handle internal traffic between services.

This piece is about where a service mesh ends and an API gateway begins, and why so many teams get that boundary wrong. They govern different traffic, not competing versions of the same traffic, and the tool you need depends entirely on which direction the request is traveling.
Why north-south and east-west traffic are fundamentally different problems
Picture two requests. One comes from someone's phone, tapping "checkout" in a shopping app, headed into a cluster from the outside world. The other comes from orders-service, already inside that same cluster, calling payments-service to confirm a charge. Both are API calls. Both move data from one place to another. That's where the similarity ends.
The phone request crosses a trust boundary. It arrives from an unknown device, on an unknown network, and someone has to check its credentials before anything downstream even wakes up. The orders-service request never leaves the building. It's a conversation between two coworkers who already know each other, and the only question is whether the conversation is fast, reliable, and private. Engineers call the first path north-south and the second east-west, and Akamai's own explainer draws that same line: a gateway handles north-south, a mesh manages east-west. Akamai even says the simplification holds up as useful, even in the cases where people argue about where exactly the edges are.
The two paths also don't scale the same way. External traffic benefits from funneling through one hardened doorway that developers and partners can point to and reason about. Internal traffic benefits from the opposite: policy applied evenly across hundreds of service pairs, with no single choke point anywhere in the middle. Mix the two up and a cluster ends up in one of two bad spots. Either internal service-to-service rules get shoved out to the edge, where they turn the gateway into a bottleneck, or external concerns get pushed into the mesh, where the clean developer-facing contract that outside partners rely on quietly falls apart. Neither failure is theoretical. Both happen the moment someone tries to make one tool do both jobs.
What an API gateway does
An API gateway is a server that sits at the front door of a system and makes sure every external request gets checked, shaped, and routed before it reaches anything internal. Its job list reads like airport security crossed with a translator: check identity, confirm the ticket, convert currency, point people to the right gate.
Start with identity. A gateway checks API keys, OAuth tokens, or JWTs on the way in, confirming who's knocking before the door even opens. It also throttles traffic, so a sudden spike from one partner, or one bad actor with a script, can't knock backend services over. It reshapes requests too, translating a REST call into gRPC or whatever format the internal services actually expect, so outside developers never need to know what's running under the hood. It manages versioning, keeping old integrations working even as the API evolves underneath them, and it runs the developer portal and documentation that let outside teams and partners actually integrate in the first place. On the reporting side, it tracks usage at the API level: who called what, how often, and from where.
None of this requires the gateway to know where every internal service instance happens to be running at any given moment. Gravitee's explanation is useful here: because backend ports and addresses change constantly, the gateway hands that lookup job off to a service-discovery layer such as Consul, Eureka, or ZooKeeper. The gateway's job is the front door, not the floor plan.
What a gateway cannot do is just as telling. It has no sidecar sitting next to every internal service, so it can't enforce mTLS between arbitrary pairs of services, can't apply circuit breaking deep inside the cluster, and can't give fine-grained control over one internal hop without touching every request that passes through the edge. Pile enough of that responsibility onto one entry point, or one scaled set of entry points, and it starts to buckle under its own job description. That's the opening a service mesh is built to fill.
What a service mesh does
A service mesh solves the problem the gateway was never built to solve: making sure every service inside the cluster can talk to every other service safely, reliably, and consistently, without engineers hand-coding that behavior into each one.
It does this by giving every service instance a small proxy that rides alongside it, built on a common proxy technology, intercepting every request that service sends or receives. MuleSoft breaks the architecture into two pieces: a data plane, made up of all those sidecar proxies handling real traffic in real time, and a control plane, a central system that defines the rules those proxies follow, from routing to security policy to monitoring.
With a proxy next to every service, the mesh can do things a single edge gateway physically cannot. It can run mutual TLS between every pair of services, so traffic is encrypted and each side can prove its identity, without a single line of code changing inside the services themselves. It can split traffic with real precision, sending a small slice of requests to a new version for a canary test or running a full blue-green rollout, all at the internal routing layer. It can apply retries, timeouts, and circuit breakers the same way across the whole cluster, so one slow service doesn't quietly take down three others downstream. It can trace a request across ten internal hops and show exactly which one is slow, instead of only reporting what happened at the edge. And it can load-balance with real awareness of which specific instance of a service is healthy right now, not just which service as a whole.
MuleSoft points out the compliance angle too: instead of wiring security and traffic rules into every service by hand, a team manages them once, in one place, which makes meeting compliance standards and keeping policy consistent across teams a lot less painful. Traffic itself still flows directly between proxies, peer to peer. The control plane only hands out configuration, never touches the actual request data, which is exactly what lets a mesh scale out horizontally instead of becoming its own bottleneck the way an overloaded gateway can.
That distributed design is the whole point. It's also the whole cost. Running a proxy next to every single service instance is a different kind of expense than running one smart front door, and that expense is where most of the real adoption debate actually lives.
Where the two tools overlap
Look at feature lists side by side and a gateway and a mesh start to look suspiciously alike. Routing happens in both. Load-balancing happens in both, too. Both produce observability data. Both enforce some form of authentication. Akamai says as much directly: both handle requests and responses, both deal with routing and service discovery, both apply rate-limiting policy. What matters is which boundary a shared capability gets applied to, not whether it exists on both lists.
Authentication is a clean example. A gateway checks a JWT or API key from a client outside the cluster. A mesh checks mTLS identity between two services already inside the cluster. Same word, authentication, completely different identity being verified. Routing works the same way: a gateway routes based on a consumer-facing URL and an API version number, while a mesh routes based on internal service identity and traffic weighting for something like a canary release. Observability splits along the same line. A gateway reports who called an API and how often. A mesh reports which internal hop between two services is slow or failing. Rate limiting follows the pattern too: gateway-side limits protect the whole cluster from outside abuse, while mesh-side limits protect one specific service from being overwhelmed by its own internal callers.
DigitalAPI's comparison states the rule in plain terms: external API security and client-facing rate limits belong at the gateway, while internal mTLS and service-to-service retries and timeouts belong in the mesh. Akamai also flags that a fair amount of the gateway-versus-mesh noise online comes from vendors marketing one tool as a replacement for the other, which muddies a distinction that's actually pretty clean once you're looking at the right axis.
That axis is the boundary being crossed, not the name of the feature. A feature called "routing" or "auth" tells you nothing about which tool should own it. Asking whether the traffic in question is crossing into the cluster or moving between services already inside it tells you everything.
The real cost of a service mesh
A service mesh is not free, and the fee gets paid in infrastructure, not dollars on an invoice. The traditional sidecar model runs a proxy process next to every single pod in the cluster, and every one of those proxies uses CPU and memory even while sitting idle. Multiply that across a large cluster and the mesh is charging the cluster a real resource toll before a single useful request has been handled.
Latency takes a hit too. Every internal service call now passes through two proxies instead of zero, one on the way out of the calling service and one on the way into the receiving service. DreamFactory frames this as a real trade-off: more granular control over traffic, paid for with an extra hop on every single internal call. This cost appears on every request, not once, for as long as the mesh is running.
The honest comparison runs the other way too. DreamFactory also points out that API gateways can hit their own scaling wall once they become the single choke point for all external traffic, while a mesh, built from the start to spread across the infrastructure, tends to scale out more gracefully. So this isn't a case where one architecture is risk-free and the other isn't: both have a place where they can bend under load, just at different points in the system.
Under the traditional sidecar model, the honest answer to "do we need a mesh" was: only if the organization is running a large number of microservices where internal reliability, observability, and security genuinely need to be enforced everywhere, consistently, without relying on every engineering team to remember to do it right. Smaller teams running a handful of services could usually get by on a gateway for external policy, handling internal concerns in application code. And the common advice to just run both a gateway and a mesh together ignored a real constraint: operating a mesh takes platform engineering capacity that not every team has sitting around. That reflects a real constraint, not a fringe opinion. What's changed recently is the size of that bill.
How ambient mesh mode changes the adoption calculus
Istio's ambient mode reached general availability in November 2024, and it rewrites the cost side of the mesh equation by getting rid of the one thing that made sidecars expensive: a dedicated proxy process sitting on every single pod.
Instead of a sidecar per pod, ambient mode runs a single node-level proxy, called a ztunnel, that handles basic layer 4 traffic (think encryption and identity) for every pod on that node at once. For teams that need more advanced layer 7 policy, things like routing based on HTTP headers or per-request authorization, ambient mode adds an optional waypoint proxy per namespace, applied only where that level of control is actually needed. Istio's own documentation lists both sidecar and ambient as supported modes today, each with guidance on when to reach for which one. The sidecar is still there, just no longer the only option on the table.
There's a real trade-off still built into ambient mode: layer 7 traffic, the kind that needs header-based routing or per-request checks, still has to pass through that waypoint proxy, which adds a hop back in. So a service that's purely doing simple internal calls sees a real latency win under ambient mode. A service leaning hard on layer 7 policy won't see the same improvement, because it's still paying for an extra hop, just a different one than before.
What this opens up is a staged path that didn't exist before. A team can roll out layer 4 mTLS and baseline observability across the entire cluster first, at a fraction of the old resource cost, and only add waypoint proxies later, for the specific services that actually need layer 7 control. That's a mesh adoption that can grow one piece at a time instead of an all-or-nothing infrastructure bet. NIST has already weighed in with a special publication covering when to use which Istio proxy mode, including the security implications of choosing ambient over sidecar. For regulated teams, this is no longer purely a performance question. For regulated teams, it's becoming a compliance one too.
How the GAMMA initiative is dissolving the gateway-mesh line
The clean north-south, east-west split this whole piece has been built on is still correct today. It's also getting less structural over time, because the specification layer underneath both tools is merging, producing a single spec that increasingly describes both kinds of traffic.
GAMMA stands for Gateway API for Mesh Management and Administration, a dedicated workstream inside the Kubernetes Gateway API subproject, created in 2022 specifically to figure out how Gateway API could describe east-west traffic, not just north-south. Its mesh-facing specifications have been part of the Gateway API Standard Channel since version 1.1.0 and are now considered generally available.
The practical effect is straightforward: a routing rule written once, as a standard HTTPRoute resource, can now be read and enforced by an ingress controller handling external traffic and by a mesh implementation handling internal traffic, using the exact same specification underneath. The two tools still do different jobs, sitting at different boundaries, solving different problems for different traffic. But the language they use to describe routing is converging into one shared vocabulary, and that's the direction the whole architecture is heading next.


