Repository navigation
Local Storage migration fails due to storage pool not found on destination KVM host #7942
Description
Activity
@rohityadavcloud
Thanks for raising the issue.
I will try to reproduce the issueI got an error when migrate the default template

and management-server log shows
2023-09-05 16:47:24,215 DEBUG [c.c.a.t.Request] (AgentManager-Handler-18:null) (logid:) Seq 2-4185251428710744296: Processing: { Ans: , MgmtId: 167780532, via: 2, Ver: v1, Flags: 110, [{"com.cloud.agent.api.MigrateAnswer":{"result":"false","details":"Exception during migrate: org.libvirt.LibvirtException: operation failed: migration of disk vda failed: Input/output error","wait":"0","bypassHostMaintenance":"false"}}] }and then
2023-09-05 16:47:24,220 DEBUG [c.c.a.t.Request] (Work-Job-Executor-4:ctx-9389653d job-33/job-34 ctx-f9253e30) (logid:baac3e91) Seq 1-1506454075355431213: Sending { Cmd , MgmtId: 167780532, via: 1(ref-trl-5592-k-Mr8-wei-zhou-kvm1), Ver: v1, Flags: 100111, [{"com.cloud.agent.api.PrepareForMigrationCommand":{"vm":{"id":"3","name":"i-2-3-VM"... ... 2023-09-05 16:47:24,815 DEBUG [c.c.a.t.Request] (AgentManager-Handler-20:null) (logid:) Seq 1-1506454075355431213: Processing: { Ans: , MgmtId: 167780532, via: 1, Ver: v1, Flags: 110, [{"com.cloud.agent.api.Answer":{"result":"false","details":"com.cloud.utils.exception.CloudRuntimeException: Could not fetch storage pool 1a30047b-99dd-4b7b-b2e5-4e862c1117a5 from libvirt due to org.libvirt.LibvirtException: Storage pool not found: no storage pool with matching uuid '1a30047b-99dd-4b7b-b2e5-4e862c1117a5' at com.cloud.hypervisor.kvm.storage.KVMStoragePoolManager.getStoragePool(KVMStoragePoolManager.java:278) at com.cloud.hypervisor.kvm.storage.KVMStoragePoolManager.getStoragePool(KVMStoragePoolManager.java:264) at com.cloud.hypervisor.kvm.storage.KVMStoragePoolManager.disconnectPhysicalDisksViaVmSpec(KVMStoragePoolManager.java:239) at com.cloud.hypervisor.kvm.resource.wrapper.LibvirtPrepareForMigrationCommandWrapper.handleRollback(LibvirtPrepareForMigrationCommandWrapper.java:150) at com.cloud.hypervisor.kvm.resource.wrapper.LibvirtPrepareForMigrationCommandWrapper.execute(LibvirtPrepareForMigrationCommandWrapper.java:62) at com.cloud.hypervisor.kvm.resource.wrapper.LibvirtPrepareForMigrationCommandWrapper.execute(LibvirtPrepareForMigrationCommandWrapper.java:52) at com.cloud.hypervisor.kvm.resource.wrapper.LibvirtRequestWrapper.execute(LibvirtRequestWrapper.java:78) at com.cloud.hypervisor.kvm.resource.LibvirtComputingResource.executeRequest(LibvirtComputingResource.java:1848) at com.cloud.agent.Agent.processRequest(Agent.java:662) at com.cloud.agent.Agent$AgentRequestHandler.doTask(Agent.java:1082) at com.cloud.utils.nio.Task.call(Task.java:83) at com.cloud.utils.nio.Task.call(Task.java:29) at java.base/java.util.concurrent.FutureTask.run(FutureTask.java:264) at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128) at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628) at java.base/java.lang.Thread.run(Thread.java:829) ","wait":"0","bypassHostMaintenance":"false"}}] }@weizhouapache I think this may be an env issue or something that affects only local storage post upgrade. If I register a new/fresh template after upgrade and then deploy VM and migrate with (local) storage, it seems to be working. The issue is the MigrationCommand fails and upon failure a prepareformigration command is sent with rollback=true.
Could you test by registering a new template again?
I suspect the I/O issue could be due to something about the migrate command not copying the template (backing file) or something else?
@weizhouapache update - I downloaded and re-registered the same template in my env, deployed a VM and that can live migrate with storage without any issues 🤦 I suspect if you were to re-register the template you were template - it might just work - now to find out why this is happening.
Could be related - #5759 cc @weizhouapache
@rohityadavcloud
Thanks for the informationThis seems to be related to #7408
@harikrishna-patnala
can you please have a look ?Update - for my specific template, it had some io_uring related settings removing those the vm migration seems to be working for newly deployed VMs. It could still be combination of other issues, not a blocker for me now. Thanks @weizhouapache
Update - for my specific template, it had some io_uring related settings removing those the vm migration seems to be working for newly deployed VMs. It could still be combination of other issues, not a blocker for me now. Thanks @weizhouapache
Thanks @rohityadavcloud for you findings.
I will investigate it tomorrow
It would be good if we can at least find a workaroundwe have found the root cause
diff --git a/plugins/hypervisors/kvm/src/main/java/com/cloud/hypervisor/kvm/resource/LibvirtConnection.java b/plugins/hypervisors/kvm/src/main/java/com/cloud/hypervisor/kvm/ +resource/LibvirtConnection.java index c70a72f399c..0f8031e3aaa 100644 --- a/plugins/hypervisors/kvm/src/main/java/com/cloud/hypervisor/kvm/resource/LibvirtConnection.java +++ b/plugins/hypervisors/kvm/src/main/java/com/cloud/hypervisor/kvm/resource/LibvirtConnection.java @@ -21,6 +21,7 @@ import java.util.Map; import org.apache.log4j.Logger; import org.libvirt.Connect; +import org.libvirt.Library; import org.libvirt.LibvirtException; import com.cloud.hypervisor.Hypervisor; @@ -44,6 +45,7 @@ public class LibvirtConnection { if (conn == null) { s_logger.info("No existing libvirtd connection found. Opening a new one"); conn = new Connect(hypervisorURI, false); + Library.initEventLoop(); s_logger.debug("Successfully connected to libvirt at: " + hypervisorURI); s_connections.put(hypervisorURI, conn); } else {@harikrishna-patnala is working on the fix.
Reacted by Rohit Yadav@weizhouapache @rohityadavcloud It seems that this is not only affecting local storage, I had the same issue that @weizhouapache reported while migrating volumes of a running VM from NFS to SharedMountPoint.
Reacted by Rohit Yadav@weizhouapache @rohityadavcloud It seems that this is not only affecting local storage, I had the same issue that @weizhouapache reported while migrating volumes of a running VM from NFS to SharedMountPoint.
@JoaoJandre
which cloudstack version do you use ?
can you provide some logs if possible ?@weizhouapache I'm using 4.19.0.0-SNAPSHOT, that I've built from main.
Unfortunately I've already rebuilt my environment with another version, but I had the same exception that you described here #7942 (comment).@weizhouapache I'm using 4.19.0.0-SNAPSHOT, that I've built from main. Unfortunately I've already rebuilt my environment with another version, but I had the same exception that you described here #7942 (comment).
Ok @JoaoJandre
If you have not seen this issue in other versions, it should be caused by same root cause which is fixed by #7945Reacted by João JandreFixed by #7945
- linked a pull request that will close this issuekvm: fix live vm migration between local storage pools #7945
on Sep 12, 2023
On a 4.18.1.0 RC2 adv zone KVM env using local only storage and local storage for systemvms, migration fails from kvm1 to kvm2 due to the following seen in the target kvm host
The above can sometimes cause the VM to crash as it fails right at the last step on the dest. kvm host end, and CloudStack may further report that domain is not running.
However, systemvm migration with storage work. I'm not sure if this is an env issue or a bug.
ISSUE TYPE
COMPONENT NAME
CLOUDSTACK VERSION
CONFIGURATION
OS / ENVIRONMENT
SUMMARY
STEPS TO REPRODUCE
EXPECTED RESULTS
ACTUAL RESULTS