Behnam Analytics

Writing Data analysis & statistics

Queueing basics for capacity planning

Little's law, why waits explode as occupancy nears 100%, why variability matters as much as volume, and Erlang C in a few lines of Python, applied to beds and clinics.

Behnam Ebrahimi 8 min read

A ward that runs at 90% occupancy on average is not 10% empty. Some days it is 80% full and some days it is full, with patients waiting in the emergency department for a bed. Average demand sets how many beds you need on average. Variation sets how many more you need to avoid a queue, and queueing theory is the arithmetic for that second part.

Four ideas cover most capacity questions an analyst gets asked: Little’s law, the utilisation curve, the cost of variability, and the Erlang C formula for pools of servers. None of them needs more than a few lines of Python.

Little’s law

For any stable system, the average number of items in it equals the arrival rate times the average time each item spends there:

L = λ × W

It holds whatever the distributions look like, as long as you measure over a long enough period. For beds, L is the average number of occupied beds, λ is admissions per day and W is the mean length of stay in days. Twenty admissions a day with a mean stay of 5.45 days gives 109 occupied beds on average.

Little’s law is the right tool for three everyday jobs:

  • Sanity-checking a plan. If a business case says admissions will grow 5% and length of stay will stay flat, occupied beds grow 5%. No model needed.
  • Converting between measures. The average number of patients waiting in the emergency department (ED) for a bed = decisions to admit per hour × mean wait in hours. If you know two, you know the third.
  • Checking a simulation or a data pipeline. In the bed occupancy simulation, the mean number of patients in a bed was 109.03, and admissions per hour times mean hours in a bed gave 109.06. If those two numbers disagree in your own data, something is being counted twice or dropped.

What Little’s law can’t tell you is how often the system is full. It is about averages. For that you need the rest of this article.

Utilisation against waiting

Take the simplest queue: one server, customers arriving at random (a Poisson process), service times that vary exponentially. This is the M/M/1 queue. Its mean wait before service, in multiples of the mean service time, is:

wait ÷ service time = ρ ÷ (1 − ρ)

where ρ is utilisation, the share of time the server is busy. At 80% utilisation, the average wait is 4 service times. At 90% it is 9. At 95% it is 19. The last five points of utilisation double the wait, and the curve heads to infinity at 100%.

Mean wait against utilisationWait in the queue, in multiples of the mean service time
Data table
Number of servers (M/M/c)
Utilisation1 server2 servers5 servers20 servers
0.51.00.30.10.0
0.511.00.30.10.0
0.521.10.40.10.0
0.531.10.40.10.0
0.541.20.40.10.0
0.551.20.40.10.0
0.561.30.50.10.0
0.571.30.50.10.0
0.581.40.50.10.0
0.591.40.50.10.0
0.61.50.60.10.0
0.611.60.60.10.0
0.621.60.60.10.0
0.631.70.70.10.0
0.641.80.70.20.0
0.651.90.70.20.0
0.661.90.80.20.0
0.672.00.80.20.0
0.682.10.90.20.0
0.692.20.90.20.0
0.72.31.00.20.0
0.712.51.00.30.0
0.722.61.10.30.0
0.732.71.10.30.0
0.742.91.20.30.0
0.753.01.30.40.0
0.763.21.40.40.0
0.773.41.50.40.0
0.783.51.60.50.1
0.793.81.70.50.1
0.84.01.80.60.1
0.814.31.90.60.1
0.824.62.00.70.1
0.834.92.20.70.1
0.845.22.40.80.1
0.855.72.60.90.1
0.866.12.81.00.1
0.876.73.11.10.2
0.887.33.41.20.2
0.898.13.81.40.2
0.99.04.31.50.3
0.9110.14.81.70.3
0.9211.55.52.00.4
0.9313.36.42.40.5
0.9415.77.62.90.6
0.9519.09.33.50.8
Service-time variability, one server (M/G/1)
UtilisationFixed length (CV 0)Low variability (CV 0.5)Exponential (CV 1)High variability (CV 1.5)
0.50.50.61.01.6
0.510.50.71.01.7
0.520.50.71.11.8
0.530.60.71.11.8
0.540.60.71.21.9
0.550.60.81.22.0
0.560.60.81.32.1
0.570.70.81.32.1
0.580.70.91.42.2
0.590.70.91.42.3
0.60.80.91.52.4
0.610.81.01.62.5
0.620.81.01.62.6
0.630.81.11.72.8
0.640.91.11.82.9
0.650.91.21.93.0
0.661.01.21.93.1
0.671.01.32.03.3
0.681.11.32.13.5
0.691.11.42.23.6
0.71.21.52.33.8
0.711.21.52.54.0
0.721.31.62.64.2
0.731.41.72.74.4
0.741.41.82.94.6
0.751.51.93.04.9
0.761.62.03.25.2
0.771.72.13.45.4
0.781.82.23.55.8
0.791.92.43.86.1
0.82.02.54.06.5
0.812.12.74.36.9
0.822.32.94.67.4
0.832.43.04.97.9
0.842.63.35.28.5
0.852.83.55.79.2
0.863.13.86.110.0
0.873.44.26.710.9
0.883.74.67.311.9
0.894.05.18.113.2
0.94.55.69.014.6
0.915.16.310.116.4
0.925.87.211.518.7
0.936.68.313.321.6
0.947.89.815.725.5
0.959.511.919.030.9

Calculated from the formulas. Source: projects/bed-occupancy-simulation. Poisson arrivals throughout. Servers: Erlang C. Variability: Pollaczek-Khinchine.

This is why “we’re only at 90%” is not reassuring, and why a small rise in demand near capacity does so much damage. In the simulation, 5% more demand raised mean bed occupancy by 3.3 points but multiplied 12-hour waits for a bed by 2.5.

The first view of the chart also shows the effect of more servers. With 20 servers sharing one queue, the wait at 90% utilisation is 0.28 service times instead of 9. Pooling is powerful, and it is the reason big hospitals can safely run hotter than small ones. More on that below.

Variability is the enemy

Switch the chart to the second view. It holds one server and Poisson arrivals, and varies only how predictable the service time is, measured by its coefficient of variation (CV, the standard deviation divided by the mean). The Pollaczek–Khinchine formula gives the mean wait exactly:

wait ÷ service time = ρ ÷ (1 − ρ) × (1 + CV²) ÷ 2

At 90% utilisation:

Service time CV Mean wait (service times)
Fixed length 0 4.50
Low variability 0.5 5.63
Exponential 1 9.00
High variability 1.5 14.63

Same utilisation, same average service time, and the wait more than triples from the most predictable service to the least. Kingman’s approximation extends this to arrivals that are not Poisson: the (1 + CV²) ÷ 2 term becomes (CV of arrivals² + CV of service²) ÷ 2. Variability in when work arrives and variability in how long it takes both cost waiting time, and they add.

For capacity planning, that means two levers besides adding capacity:

  • Smooth what you control. Emergency demand is random, but elective admissions, clinic templates and discharge timing are scheduled by the organisation. Spreading them evenly across the week removes variability that nobody had to create.
  • Reduce the spread of the work. A clinic where some appointments take 10 minutes and others 40 queues worse than one with the same mean and a narrower spread. Booking long and short appointments into separate slots is a variability fix, not an efficiency one.

Pools of servers: M/M/c and Erlang C

Beds, clinic rooms and call handlers are pools: c identical servers sharing one queue. The offered load, a, is the arrival rate times the mean service time, measured in servers. For beds, that is admissions per day times mean stay in days, the L from Little’s law. Utilisation is a ÷ c.

The Erlang C formula gives the probability that an arrival has to wait. The direct formula has factorials that overflow for large pools, so I compute it through the Erlang B recursion, which is stable for any size:

def erlang_b(servers: int, load: float) -> float:
    """Chance that all servers are busy in a loss system, by the stable recursion."""
    blocked = 1.0
    for k in range(1, servers + 1):
        blocked = load * blocked / (k + load * blocked)
    return blocked


def erlang_c(servers: int, load: float) -> float:
    """Chance that an arrival has to wait in an M/M/c queue.

    load is the offered load in servers: arrival rate x mean service time. For beds, that is
    admissions per day x mean length of stay in days.
    """
    utilisation = load / servers
    if utilisation >= 1:
        return 1.0
    blocked = erlang_b(servers, load)
    return blocked / (1 - utilisation * (1 - blocked))


def mean_wait(servers: int, load: float, mean_service: float) -> float:
    """Mean wait in the queue for an M/M/c system, in the units of mean_service."""
    return erlang_c(servers, load) * mean_service / (servers - load)

A worked example: 120 beds, 20 emergency admissions a day, a mean stay of 5.45 days. The offered load is 109.0 beds, and utilisation is 90.8%.

Case Utilisation Chance of waiting Mean wait
120 beds 90.8% 21.6% 2.57 hours
125 beds 87.2% 8.9% 0.73 hours
120 beds, 5% more demand 95.4% 50.1% 11.81 hours
125 beds, 5% more demand 91.6% 24.2% 3.00 hours

Five more beds cut the chance of waiting from 21.6% to 8.9%. Five per cent more demand on the same 120 beds pushes it to 50.1% and the mean wait to almost 12 hours. To keep the chance of waiting at or below 10%, this unit needs 125 beds today and 131 with 5% more demand. Those are the numbers a business case for beds should contain, not “we need 92% occupancy”.

Size changes the safe occupancy

The same calculation shows why a single occupancy target can’t suit every unit. Here is the highest utilisation each pool size can run at while keeping the chance of waiting at or below 10%:

Servers (beds) Highest utilisation
10 59.9%
30 75.6%
60 82.4%
120 87.4%
250 91.2%
500 93.7%
Chance a patient has to wait for a bed, by size of bed poolErlang C probability of waiting at each average occupancy
Data table
Average occupancy10 beds30 beds120 beds500 beds
0.610%1%0%0%
0.6111%1%0%0%
0.6212%1%0%0%
0.6313%1%0%0%
0.6414%2%0%0%
0.6515%2%0%0%
0.6617%2%0%0%
0.6718%3%0%0%
0.6819%3%0%0%
0.6921%4%0%0%
0.722%4%0%0%
0.7124%5%0%0%
0.7225%6%0%0%
0.7327%7%0%0%
0.7429%8%0%0%
0.7531%9%0%0%
0.7633%10%0%0%
0.7734%12%0%0%
0.7837%14%0%0%
0.7939%15%1%0%
0.841%17%1%0%
0.8143%19%2%0%
0.8246%22%2%0%
0.8348%24%3%0%
0.8450%27%4%0%
0.8553%30%5%0%
0.8656%33%7%0%
0.8758%36%9%0%
0.8861%40%12%0%
0.8964%43%14%1%
0.967%47%18%1%
0.9170%51%22%2%
0.9273%56%27%4%
0.9376%60%33%7%
0.9479%65%40%11%
0.9583%70%47%18%
0.9686%76%56%27%
0.9789%82%65%39%
0.9893%87%76%55%
0.9996%94%87%75%

Calculated from the formulas. Source: projects/bed-occupancy-simulation. Assumes steady Poisson arrivals and exponential stays, which flatters real wards.

A 10-bed unit at 85% occupancy is in far more trouble than a 500-bed hospital at 85%. The spare capacity needed grows roughly with the square root of the load, not in proportion to it, so small units need proportionally more slack. This is also the case for pooling: two 30-bed wards that can take each other’s patients behave more like one 60-bed pool. The catch is “can”: a pooled bed only helps if the patient can safely use it.

What this means for bed and clinic capacity

  • Plan beds from the distribution, not the average. Little’s law gives the average. The spare capacity comes from queueing, and it depends on how much you are willing to let patients wait.
  • Set occupancy expectations by pool size. A 92% target means something different for a 12-bed specialist ward and for a whole hospital. Bagust and colleagues’ 1999 BMJ simulation found risks becoming discernible above about 85% average occupancy, and regular bed shortages at 90% or more, for a whole acute hospital. The table above says smaller units hit trouble sooner.
  • Expect non-linear responses near capacity. A 5% rise in demand is not a 5% rise in waits. Report both, so nobody is surprised.
  • Remove the variability you create. Uneven elective scheduling, one daily ward round that releases every discharge at 16:00, and clinics booked without regard to appointment length all add waiting that capacity then has to absorb.
  • For clinics and waiting lists, the same shape applies. A service whose weekly capacity only just exceeds average weekly demand will see its waiting list swing widely and recover slowly after every bad week. The closer to 100% utilisation, the longer each recovery takes.

Where the formulas stop

Erlang C assumes arrivals at a constant random rate, exponential service times, first-come-first-served, and nothing that depends on the clock. Real beds break all of that. Admissions follow daily and weekly cycles, discharges bunch in the afternoon, weekends discharge fewer patients, and a crowded hospital behaves differently from a quiet one.

So I use the formulas for orientation and simulation for decisions. With inputs that match the worked example (20 admissions a day, 5.45 days in a bed on average), the bed occupancy simulation, which adds hourly and weekly cycles, lognormal stays and weekend discharge delays, gave a chance of waiting of 21.0% against Erlang C’s 21.6%, but a mean wait of 1.86 hours against 2.57. With 5% more demand, the gap widened: 37.5% against 50.1%, and 4.03 hours against 11.81. Part of that gap is a choice in the simulation, where time waiting in the ED counts towards the stay, so a crowded unit gets some bed-days back. Erlang C has no such feedback. At today’s load, the formula was close on how often patients wait and too pessimistic on how long, and it drifted further off as the load rose.

That is the usual pattern. The formula tells you which way things move and roughly how fast. Questions about timing, such as discharging before noon, need a model with a clock in it. Simulation for what-if questions covers how to build and trust one.

Reproduce

The formulas are in projects/bed-occupancy-simulation/queueing.py. The charts and every number in this article come from:

uv run python projects/bed-occupancy-simulation/run.py