Live migration of virtual machines
Christopher ClarkK. FraserS. HandJ. G. HansenE. JulC. LimpachI. PrattA. Warfield
Demonstrates an iterative pre-copy live migration system implemented in Xen that dynamically adapts network bandwidth to transfer running virtual machines with service downtimes as low as tens of milliseconds while preserving active network connections.
Researchers developed a system for live migration of entire operating system instances between physical hosts in data centers and clusters using the Xen virtual machine monitor. The work addressed the need for clean separation of hardware and software to support fault management, load balancing, and maintenance without residual dependencies on the original machine or disruption to running services.
The project set out to evaluate a pre-copy approach that iteratively transfers memory pages while the virtual machine continues executing, followed by a short stop-and-copy phase, and to measure its impact on downtime and total migration time across realistic workloads. The team implemented the system on commodity hardware, tracked page dirtying with shadow page tables, applied dynamic network rate limiting to control contention, and tested it on benchmarks including a static web server, SPECweb99, an OLTP database, a Quake 3 server, and a synthetic high-dirtying workload.
The analysis showed that live migration is practical for most server loads. Downtime reached as low as 60 ms for the Quake 3 server and 210 ms for a heavily loaded SPECweb99 instance, with total migration times of roughly one minute and negligible effects on client-visible performance. Pre-copy reduced downtime by factors of four to sixteen compared with naive stop-and-copy, and the writable working set of typical workloads proved small enough to allow effective iteration. Extremely high dirtying rates remained rare and produced longer but still bounded outages.
These results mean administrators can move live virtual machines for hardware servicing or load balancing with almost no service interruption and without requiring guest operating system cooperation or network changes beyond a simple ARP update. The approach therefore strengthens virtualization as a management tool in clusters while avoiding the fragility of process-level migration.
Further development is needed to handle wide-area networks, local disk migration, and automated cluster-wide placement decisions that prioritize low writable working set instances first. Current limitations include the assumption that the source host remains stable until migration commits and reduced effectiveness for workloads that continuously dirty memory faster than available network bandwidth. Overall confidence in the core findings is high because the evaluation covered multiple representative workloads on real hardware with consistent quantitative results.
- Paper: Xen and the art of virtualization, P. Barham et al. (2003). Read Xen’s foundational virtualization design first: the migration system is implemented in Xen and relies on its virtual-machine architecture.
No sufficiently relevant recommendations were found.
