Live migration of virtual machines

Christopher ClarkK. FraserS. HandJ. G. HansenE. JulC. LimpachI. PrattA. Warfield

article2005NSDI3,264 citations

Demonstrates an iterative pre-copy live migration system implemented in Xen that dynamically adapts network bandwidth to transfer running virtual machines with service downtimes as low as tens of milliseconds while preserving active network connections.

Listen

Researchers developed a system for live migration of entire operating system instances between physical hosts in data centers and clusters using the Xen virtual machine monitor. The work addressed the need for clean separation of hardware and software to support fault management, load balancing, and maintenance without residual dependencies on the original machine or disruption to running services.

The project set out to evaluate a pre-copy approach that iteratively transfers memory pages while the virtual machine continues executing, followed by a short stop-and-copy phase, and to measure its impact on downtime and total migration time across realistic workloads. The team implemented the system on commodity hardware, tracked page dirtying with shadow page tables, applied dynamic network rate limiting to control contention, and tested it on benchmarks including a static web server, SPECweb99, an OLTP database, a Quake 3 server, and a synthetic high-dirtying workload.

The analysis showed that live migration is practical for most server loads. Downtime reached as low as 60 ms for the Quake 3 server and 210 ms for a heavily loaded SPECweb99 instance, with total migration times of roughly one minute and negligible effects on client-visible performance. Pre-copy reduced downtime by factors of four to sixteen compared with naive stop-and-copy, and the writable working set of typical workloads proved small enough to allow effective iteration. Extremely high dirtying rates remained rare and produced longer but still bounded outages.

These results mean administrators can move live virtual machines for hardware servicing or load balancing with almost no service interruption and without requiring guest operating system cooperation or network changes beyond a simple ARP update. The approach therefore strengthens virtualization as a management tool in clusters while avoiding the fragility of process-level migration.

Further development is needed to handle wide-area networks, local disk migration, and automated cluster-wide placement decisions that prioritize low writable working set instances first. Current limitations include the assumption that the source host remains stable until migration commits and reduced effectiveness for workloads that continuously dirty memory faster than available network bandwidth. Overall confidence in the core findings is high because the evaluation covered multiple representative workloads on real hardware with consistent quantitative results.

  • Paper: Xen and the art of virtualization, P. Barham et al. (2003). Read Xen’s foundational virtualization design first: the migration system is implemented in Xen and relies on its virtual-machine architecture.

No sufficiently relevant recommendations were found.

Cover for Live migration of virtual machines

Abstract

Migrating operating system instances across distinct physical hosts is a useful tool for administrators of data centers and clusters: It allows a clean separation between hardware and software, and facilitates fault management, load balancing, and low-level system maintenance.

By carrying out the majority of migration while OSes continue to run, we achieve impressive performance with minimal service downtimes; we demonstrate the migration of entire OS instances on a commodity cluster, recording service downtimes as low as 60ms. We show that that our performance is sufficient to make live migration a practical tool even for servers running interactive loads.

In this paper we consider the design options for migrating OSes running services with liveness constraints, focusing on data center and cluster environments. We introduce and analyze the concept of writable working set, and present the design, implementation and evaluation of high-performance OS migration built on top of the Xen VMM.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Design
  • 3.1 Migrating Memory
  • 3.2 Local Resources
  • 3.3 Design Overview
  • Tracking the Writable Working Set of SPEC CINT2000
  • 4 Writable Working Sets
  • 4.1 Measuring Writable Working Sets
  • 4.2 Estimating Migration Effectiveness
  • 5 Implementation Issues
  • 5.1 Managed Migration
  • 5.2 Self Migration
  • 5.3 Dynamic Rate-Limiting
  • 5.4 Rapid Page Dirtying
  • 5.5 Paravirtualized Optimizations
  • 6 Evaluation
  • 6.1 Test Setup
  • 6.2 Simple Web Server
  • 6.3 Complex Web Workload: SPECweb99
  • 6.4 Low-Latency Server: Quake 3
  • 6.5 A Diabolical Workload: MMuncher
  • 7 Future Work
  • 7.1 Cluster Management
  • 7.2 Wide Area Network Redirection
  • 7.3 Migrating Block Devices
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — Transactional Pre-Copy Live Virtual Machine Migration Lifecycle

    model/method

    Live migration of a virtual machine (VM) running an operating system and active network services proceeds through a six-stage transactional protocol between a source physical host AA and a destination physical host BB:

    1. Stage 0 (Pre-Migration): An active VM runs on host AA. A target host BB may be preselected, and shared storage resources (such as network-attached storage) are maintained.
    2. Stage 1 (Reservation): A migration request is issued from AA to BB. Host BB confirms that sufficient CPU and memory resources exist and reserves a VM container of identical size. If resource allocation fails, the VM continues execution on AA unaffected.
    3. Stage 2 (Iterative Pre-Copy): While the VM continues normal execution on AA, memory pages are iteratively copied across the network to BB. In round 0, the entire physical memory allocation of the VM is transferred. In each subsequent round nn, only pages dirtied during round n−1n-1 are sent.
    4. Stage 3 (Stop-and-Copy): The running VM on host AA is suspended. Network traffic redirection is initiated. Checkpointed virtual CPU (VCPU) register state, device state, and any remaining dirty memory pages are transferred to BB. At the end of this stage, identical, consistent copies of the VM exist on both hosts, with AA remaining the primary authority.
    5. Stage 4 (Commitment): Host BB signals successful receipt of the complete consistent image to host AA. Host AA acknowledges this message, committing the transaction. Host AA discards the local VM instance, and host BB becomes the primary host.
    6. Stage 5 (Activation): The VM on host BB is activated and resumes execution from its checkpointed state. Post-migration code reconnects virtual device drivers to local backends and advertises its network location via unsolicited ARP.

    If any hardware or network failure occurs prior to the Stage 4 commitment acknowledgment, the migration transaction aborts and the VM safely resumes local execution on host AA.

  2. Knowl 2 — Dynamic Network Rate-Limiting for Iterative Pre-Copy Migration

    algorithm

    To balance network contention with active services against final service downtime, network transmission bandwidth is dynamically adjusted between pre-copy rounds.

    Input: Minimum bandwidth limit BminB_{min}, maximum bandwidth limit BmaxB_{max}, constant rate increment ΔB=50 Mbit/s\Delta B = 50\text{ Mbit/s}, remaining memory threshold Mthresh=256 KBM_{thresh} = 256\text{ KB}, total memory pages NtotalN_{total}
    Output: Iterative pre-copy memory transfer to destination host
    Round n=0n = 0
    Bandwidth limit B0=BminB_0 = B_{min}
    Transfer all NtotalN_{total} memory pages from source to destination at rate B0B_0
    Record duration of round T0T_0
    loop:
        n=n+1n = n + 1
        Count pages dirtied during round n−1n-1, denoted Dn−1D_{n-1}
        Mremain=Dn−1×page_sizeM_{remain} = D_{n-1} \times \text{page\_size}
        DirtyRaten−1=Mremain/Tn−1\text{DirtyRate}_{n-1} = M_{remain} / T_{n-1}
        Bn=DirtyRaten−1+ΔBB_n = \text{DirtyRate}_{n-1} + \Delta B
        
        if Bn>BmaxB_n > B_{max} or Mremain<MthreshM_{remain} < M_{thresh}:
            break loop
            
        Transfer Dn−1D_{n-1} dirtied memory pages to destination at rate BnB_n
        Record duration of round TnT_n
    Suspend VM on source host
    Transfer all remaining dirty pages at maximum rate BmaxB_{max}
    Resume VM on destination host
  3. Knowl 3 — Dirty Page Tracking via Xen Shadow Page Tables

    model/method

    In paravirtualized Xen environments, guest operating systems normally manage their own hardware-walked page tables that map virtual addresses directly to machine physical addresses. To perform live memory migration without modifying or halting the guest OS, Xen interposes shadow page tables:

    1. Read-Only Shadowing: Xen constructs shadow page tables that mirror guest translations but mark all page table entries (PTEs) as read-only, irrespective of permissions granted in the guest page tables. The hardware memory management unit (MMU) is configured to walk these shadow page tables.
    2. Fault Trapping and Dirty Logging: When the guest operating system writes to a page, a write page fault is generated and intercepted by Xen. Xen verifies whether the guest PTE permits write access. If permitted, Xen sets the corresponding bit in a per-VM physical dirty bitmap and updates the shadow PTE with write permission, allowing subsequent writes without hypervisor trapping.
    3. Iteration Advancement: At the start of each pre-copy round, the dirty bitmap is copied to the migration daemon. Xen's internal dirty bitmap is then cleared, and all shadow page tables are destroyed and recreated with read-only permissions. Consequently, subsequent memory modifications are tracked for the next round.
  4. Knowl 4 — Writable Working Set in Virtual Machine Migration

    definition

    The writable working set (WWS) of an operating system or virtual machine is the subset of memory pages that are modified with high frequency during workload execution.

    In iterative pre-copy migration, pages that are not part of the WWS need to be transmitted across the network only once during early iterations. In contrast, pages belonging to the WWS (such as execution stacks, local process variables, and active network/disk I/O buffers) are repeatedly modified faster than, or at rates comparable to, network transfer speeds. Pre-copying WWS pages during iterative rounds wastes network and CPU bandwidth; thus, the size of the WWS establishes the lower bound on the volume of dirty memory that must be transferred during the final stop-and-copy suspension phase.

  5. Knowl 5 — Layer-2 Network Connection Preservation via Unsolicited ARP

    model/method

    To preserve open network connections (such as active TCP connections) without relying on forwarding daemons on the source host or client-side reconnection protocols, live migration leverages standard switched local area network (LAN) infrastructure:

    1. State Preservation: The virtual machine migrates its complete network stack state, including IP addresses, TCP Protocol Control Blocks (PCBs), socket queues, and sequence numbers.
    2. Gratuitous Address Resolution: Immediately upon resuming execution on destination host BB, the VM or hypervisor issues an unsolicited broadcast ARP (Address Resolution Protocol) reply. This packet announces that the migrating IP address is now bound to the network interface MAC address of host BB.
    3. Network Switch Relearning: Network switches update their forwarding tables to route subsequent IP packets directly to the new physical switch port on host BB. For environments where broadcast ARP is filtered by routers, the guest OS sends directed ARP replies directly to all peers present in its local ARP cache, or retains its original virtual MAC address and allows the Ethernet switch learning mechanism to update port associations.
  6. Knowl 6 — Dirty Page Scanning Optimizations: Bitmap Peeking and Pseudo-Random Traversal

    model/method

    To optimize the selection of pages transferred during iterative pre-copy, two scanning heuristics are applied:

    • Dirty Bitmap Peeking: During pre-copy iteration round nn, the transfer daemon sequentially iterates through the dirty list generated in round n−1n-1. Before transmitting a page, the daemon inspects ('peeks' at) the active dirty bitmap being filled concurrently by Xen for round nn. If the page has already been redirtied in round nn, its transmission in the current round is aborted and deferred, avoiding the wasted network transmission of a stale page.
    • Pseudo-Random Memory Scanning: Because memory modifications exhibit physical clustering (where dirtied pages are spatially adjacent to other dirtied pages), sequential memory scanning can repeatedly encounter clustered write bursts and miss redirtied pages. Scanning the physical memory address space in a pseudo-random permutation mitigates spatial dirtying correlation and improves the filtering efficiency of bitmap peeking.
  7. Knowl 7 — Guest OS Self-Migration Architecture and Two-Stage Shadow Buffering

    model/method

    As an alternative to hypervisor-managed migration, self-migration embeds the migration logic directly inside the guest operating system kernel, requiring no modifications to the hypervisor on the source machine:

    • In-Guest Dirty Tracking: At the beginning of each round, the guest kernel marks all virtual address space mappings as read-only. A designated spare bit inside each page table entry (PTE) is reserved to distinguish migration-tracking write faults from standard application page faults (e.g., copy-on-write). When a write fault occurs on a migration-protected page, the kernel logs the physical page index into a dirty bitmap and restores write permissions.
    • Two-Stage Stop-and-Copy with Shadow Buffering: Because the migrating OS must remain active to execute the transfer code during the final phase, consistent checkpointing is achieved via a two-stage process:
      1. Stage 1: All non-migration operating system activity and interrupts are suspended. The kernel performs a final scan of the dirty bitmap, transmitting dirty pages across the network and clearing each dirty bit upon transfer. Any page dirtied by the migration thread itself during this scan is copied into an in-kernel shadow buffer.
      2. Stage 2: The contents of the shadow buffer are transferred over the network to the destination machine. Updates occurring during this brief final stage are ignored.
  8. Knowl 8 — SPECweb99 Live Migration Performance and Service Downtime

    empirical result

    Live migration was evaluated on an 800MB virtual machine running XenLinux 2.4.27 and Apache 1.3 hosting the SPECweb99 benchmark (a mixed workload comprising 30% dynamic content generation, 16% HTTP POSTs, and 0.5% CGI scripts) under 90% overload capacity (350 conformant client connections, where conformance requires aggregate bandwidth >320 Kbit/s> 320\text{ Kbit/s} per user) across Dell PE-2650 dual Xeon 2GHz hosts on Gigabit Ethernet:

    • Bandwidth Progression: The initial pre-copy round operated at a low rate, taking 54.1 s54.1\text{ s} to transfer 676.8 MB676.8\text{ MB}. Subsequent rounds progressively adapted transfer bandwidth, transferring 126.7 MB126.7\text{ MB}, 39.0 MB39.0\text{ MB}, 28.4 MB28.4\text{ MB}, 24.2 MB24.2\text{ MB}, 16.7 MB16.7\text{ MB}, 14.2 MB14.2\text{ MB}, 15.3 MB15.3\text{ MB}, and leaving 18.2 MB18.2\text{ MB} for the final round.
    • Total Migration Time and Data: Total migration time was 71 s71\text{ s}, and total transferred data was 960 MB960\text{ MB} (1.20×1.20\times the virtual machine's allocated physical memory).
    • Downtime and Conformance: The final stop-and-copy suspension required 201 ms201\text{ ms} to transfer the remaining 18.2 MB18.2\text{ MB}, plus 9 ms9\text{ ms} of VM startup overhead, yielding a total service downtime of 210 ms210\text{ ms}. All 350 client connections remained conformant with zero session drops.
  9. Knowl 9 — Quake 3 Server Live Migration Latency and Downtime

    empirical result

    Live migration was evaluated on a 64MB virtual machine hosting an interactive multi-player Quake 3 game server with 6 active players in a shared arena over Gigabit Ethernet:

    • Data Transferred: Total migration transmitted 88 MB88\text{ MB} of data (1.37×1.37\times total memory allocation) over successive rate-adapted rounds (56.3 MB56.3\text{ MB}, 20.4 MB20.4\text{ MB}, 4.6 MB4.6\text{ MB}, 1.6 MB1.6\text{ MB}, 1.2 MB1.2\text{ MB}, 0.9 MB0.9\text{ MB}, 1.2 MB1.2\text{ MB}, 1.1 MB1.1\text{ MB}, 0.8 MB0.8\text{ MB}, 0.2 MB0.2\text{ MB}, and 0.1 MB0.1\text{ MB}).
    • Downtime: The final stop-and-copy phase transferred 148 KB148\text{ KB} of remaining data in 20 ms20\text{ ms}, with 40 ms40\text{ ms} spent on startup overhead, resulting in a total service outage of 60 ms60\text{ ms} (and 48 ms48\text{ ms} to 50 ms50\text{ ms} in repeated test runs).
    • User Latency: Network packet inter-arrival flight times measured at client consoles exhibited a transient latency increase of ≈50 ms\approx 50\text{ ms}, causing no perceptible disruption to active gameplay.
  10. Knowl 10 — Paravirtualized Migration Optimizations: Process Stunning and Memory Ballooning

    model/method

    When guest operating systems are paravirtualized to participate cooperatively in live migration, two kernel-level optimizations can be utilized:

    • Rogue Process Stunning: Synthetic or pathological applications can dirty memory at rates exceeding network line speeds (e.g., dirtying memory at up to 320 Gbit/s320\text{ Gbit/s} by modifying one word per page). An in-kernel monitoring thread measures the per-process write fault rate during migration iterations. When a process exceeds a threshold of dirtying activity (e.g., 40 write faults within an iteration), the kernel stuns the process by placing it onto a wait queue until the migration completes, preventing unbounded pre-copy iterations.
    • Page Cache Ballooning/Freeing: Prior to the first iterative pre-copy pass, the guest OS kernel can identify cold buffer cache pages and free memory pages and release them back to the hypervisor via memory ballooning. This eliminates unneeded memory pages from the initial full memory scan, reducing the initial iteration duration and overall data transmission volume.
  11. Knowl 11 — Environmental Assumptions and Workload Limitations of Pre-Copy Live Migration

    limitation

    The pre-copy live migration architecture relies on specific environment and workload assumptions:

    • Local Switched Subnet Requirement: The unsolicited ARP mechanism relies on the source and destination physical hosts residing on the same layer-2 broadcast domain / IP subnet. Migrating across wide-area networks (WANs) or distinct routing domains requires additional network tunneling or mobile indirection layers.
    • Shared Storage Dependency: The system assumes storage is decoupled from the host and hosted on shared network-attached storage (NAS) or block devices (e.g., iSCSI or GNBD). Live migration of local physical disk blocks is not supported natively.
    • Pathological Memory Dirtying (Diabolical Workload): In workloads where active memory write throughput continuously exceeds network bandwidth across the entire working set (evaluated using an 'MMuncher' test program continuously writing to a 256 MB256\text{ MB} memory region in a 512 MB512\text{ MB} VM), pre-copy iterations cannot converge. The dynamic rate-limiting algorithm escalates transmission bandwidth up to maximum capacity (500 Mbit/s500\text{ Mbit/s}), detects lack of convergence, and terminates pre-copying, forcing a full stop-and-copy suspension phase that results in an extended service downtime (3.5 s3.5\text{ s}).

Coverage note — Omitted qualitative background surveys of 1980s process migration systems (Sprite, MOSIX, DEMOS/MP) and preliminary speculative future work regarding wide-area DNS redirection and RAID-5 local disk mirroring, as they do not constitute validated core contributions of the paper.

References

  1. 1.Paul Barham, Boris Dragovic, Keir Fraser, Steven Hand, Tim Harris, Alex Ho, Rolf Neugebauer, Ian Pratt, and Andrew Warfield. Xen and the art of virtualization. In Proceedings of the nineteenth ACM symposium on Operating Systems Principles (SOSP19), pages 164–177. ACM Press, 2003.
  2. 2.D. Milojicic, F. Douglis, Y. Paindaveine, R. Wheeler, and S. Zhou. Process migration. ACM Computing Surveys, 32(3):241–299, 2000.
  3. 3.C. P. Sapuntzakis, R. Chandra, B. Pfaff, J. Chow, M. S. Lam, and M.Rosenblum. Optimizing the migration of virtual computers. In Proc. of the 5th Symposium on Operating Systems Design and Implementation (OSDI-02), December 2002.
  4. 4.M. Kozuch and M. Satyanarayanan. Internet suspend/resume. In Proceedings of the IEEE Workshop on Mobile Computing Systems and Applications, 2002.
  5. 5.Andrew Whitaker, Richard S. Cox, Marianne Shaw, and Steven D. Gribble. Constructing services with interposable virtual hardware. In Proceedings of the First Symposium on Networked Systems Design and Implementation (NSDI ’04), 2004.
  6. 6.S. Osman, D. Subhraveti, G. Su, and J. Nieh. The design and implementation of zap: A system for migrating computing environments. In Proc. 5th USENIX Symposium on Operating Systems Design and Implementation (OSDI-02), pages 361–376, December 2002.
  7. 7.Jacob G. Hansen and Asger K. Henriksen. Nomadic operating systems. Master’s thesis, Dept. of Computer Science, University of Copenhagen, Denmark, 2002.
  8. 8.Hermann Härtig, Michael Hohmuth, Jochen Liedtke, and Sebastian Schönberg. The performance of microkernel-based systems. In Proceedings of the sixteenth ACM Symposium on Operating System Principles, pages 66–77. ACM Press, 1997.
  9. 9.VMWare, Inc. VMWare VirtualCenter Version 1.2 User’s Manual. 2004.
  10. 10.Michael L. Powell and Barton P. Miller. Process migration in DEMOS/MP. In Proceedings of the ninth ACM Symposium on Operating System Principles, pages 110–119. ACM Press, 1983.
  11. 11.Marvin M. Theimer, Keith A. Lantz, and David R. Cheriton. Preemptable remote execution facilities for the V-system. In Proceedings of the tenth ACM Symposium on Operating System Principles, pages 2–12. ACM Press, 1985.
  12. 12.Eric Jul, Henry Levy, Norman Hutchinson, and Andrew Black. Fine-grained mobility in the emerald system. ACM Trans. Comput. Syst., 6(1):109–133, 1988.
  13. 13.Fred Douglis and John K. Ousterhout. Transparent process migration: Design alternatives and the Sprite implementation. Software - Practice and Experience, 21(8):757–785, 1991.
  14. 14.A. Barak and O. La’adan. The MOSIX multicomputer operating system for high performance cluster computing. Journal of Future Generation Computer Systems, 13(4-5):361–372, March 1998.
  15. 15.J. K. Ousterhout, A. R. Cherenson, F. Douglis, M. N. Nelson, and B. B. Welch. The Sprite network operating system. Computer Magazine of the Computer Group News of the IEEE Computer Group Society, ; ACM CR 8905-0314, 21(2), 1988.
  16. 16.E. Zayas. Attacking the process migration bottleneck. In Proceedings of the eleventh ACM Symposium on Operating systems principles, pages 13–24. ACM Press, 1987.
  17. 17.Peter J. Denning. Working Sets Past and Present. IEEE Transactions on Software Engineering, SE-6(1):64–84, January 1980.
  18. 18.Jacob G. Hansen and Eric Jul. Self-migration of operating systems. In Proceedings of the 11th ACM SIGOPS European Workshop (EW 2004), pages 126–130, 2004.
  19. 19.C. E. Perkins and A. Myles. Mobile IP. Proceedings of International Telecommunications Symposium, pages 415–419, 1997.
  20. 20.Alex C. Snoeren and Hari Balakrishnan. An end-to-end approach to host mobility. In Proceedings of the 6th annual international conference on Mobile computing and networking, pages 155–166. ACM Press, 2000.

Citation

MLA
Clark, C. J., et al. “Live Migration of Virtual Machines”. Research at the University of Copenhagen (University of Copenhagen), 2005, pp. 273–86, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.120.374.
APA
Clark, C. J., Fraser, K., Hand, S., Hansen, J. G., Jul, E., Limpach, C., Pratt, I. A., & Warfield, A. (2005). Live migration of virtual machines. Research at the University of Copenhagen (University of Copenhagen), 273–286. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.120.374
Chicago
Clark, C. J., K. Fraser, S. Hand, et al. 2005. “Live Migration of Virtual Machines”. Research at the University of Copenhagen (University of Copenhagen), 273–86. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.120.374.
Harvard
Clark, C.J. et al. (2005) “Live migration of virtual machines”, Research at the University of Copenhagen (University of Copenhagen), pp. 273–286. Available at: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.120.374.
Vancouver
1. Clark CJ, Fraser K, Hand S, Hansen JG, Jul E, Limpach C, Pratt IA, Warfield A (2005) Live migration of virtual machines. Research at the University of Copenhagen (University of Copenhagen) 273–286

BibTeX

@article{clark2005live,
  title = {Live migration of virtual machines},
  author = {Clark, Christopher J. and Fraser, Keir and Hand, Steven and Hansen, Jacob Gorm and Jul, Eric and Limpach, Christian and Pratt, Ian A. and Warfield, Andrew},
  year = {2005},
  journal = {Research at the University of Copenhagen (University of Copenhagen)},
  pages = {273-286},
  url = {http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.120.374}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF