For the shared families NX and NP, the hypervisor schedules multiple instances on the same physical core. Each vCPU gets a quota; if you do not use it, you give your share to your neighbours. This is why these families are cheap: you pay for an expected value, not a guarantee.
If a neighbour computes while you want to compute, your instance waits. This waiting time is visible in the guest system — as %st in top, vmstat or mpstat. Steal time is not a fault, but the honest display of a shared system. On NX and NP, we measure a monthly average of under 3%, and up to 12% during load peaks on individual hosts. On dedicated families, it is zero by design.
We regulate burst behaviour via a credit model instead of hard throttling: a shared vCPU can draw full core performance for up to 20 minutes per hour. After that, the base share applies — 25% per vCPU for NX, 40% for NP. Unused credit is saved for up to eight hours. A build server that compiles for ten minutes and waits for fifty minutes never sees this limit. A video transcoder sees it after twenty minutes.
The rule of thumb from our operations: if the average load of a vCPU is permanently above 20% over a week, a dedicated instance is not only faster, but also cheaper per unit of work done.