Repository navigation
PBSControllerLauncher: Unable to connect_client_sync() #619
Description
Activity
The program hangs in the while loop at
because the config files created inipyparallel/ipyparallel/cluster/launcher.py
Line 339 in 527d7b2
while not all(os.path.isfile(f) for f in paths): PROFILE/securityare namedipcontroller-client.jsonandipcontroller-engine.jsonwhereas the program expects the filenames to includecluster_id, e.g.ipcontroller-1635416202-ou8n-client.json,ipcontroller-1635416202-ou8n-engine.json. Is this also fixed in #606?This is why I need to get CI tests for all the non-slurm batch launchers (#604)!
I do believe the issue is fixed in dev, but those custom templates will still reintroduce the problem. If you add
--cluster-id={cluster_id}it will be fixed. The generic fix is to use{program_and_args}instead ofipengine --profile-dir={profile_dir}.I believe these templates will work:
c.PBSControllerLauncher.batch_template = ''' #PBS -N ipcontroller #PBS -V #PBS -j oe #PBS -l walltime=01:00:00 #PBS -l nodes=1:ppn=1 cd $PBS_O_WORKDIR conda activate ipp7 {program_and_args} ''' c.PBSEngineSetLauncher.batch_template = ''' #PBS -N ipengine #PBS -j oe #PBS -V #PBS -l walltime=01:00:00 #PBS -l nodes={n//4}:ppn=4 cd $PBS_O_WORKDIR conda activate ipp7 module load intel mpiexec -n {n} {program_and_args} '''
The next release uses environment variables to pass things like the cluster id, which means you must add
#PBS -Vto your options.Reacted by Lukas KoschmiederThank you for the quick reply! The general method works. 👍
Is it possible to instantiate a
Clusterwithout an existing IPython profile andipcluster_config.pyby passingc.PBSControllerLauncher.batch_templateandc.PBSEngineSetLauncher.batch_templatesomehow directly to the class constructor from a Jupyter Notebook?Pseudocode:
controller_template=''' #PBS -N ipcontroller #PBS -j oe #PBS -l walltime=01:00:00 #PBS -l nodes=1:ppn=1 ##PBS -q {queue} cd $PBS_O_WORKDIR conda activate ipp7 {program_and_args} ''' engine_template = ''' #PBS -N ipengine #PBS -j oe #PBS -l walltime=01:00:00 #PBS -l nodes={n//4}:ppn=4 ##PBS -q {queue} cd $PBS_O_WORKDIR conda activate ipp7 module load intel mpiexec -n {n} {program_and_args} ''' cluster=ipp.Cluster( n=4, controller_ip='*', profile='pbs-2021-10-28', extra_options={ 'c.PBSControllerLauncher.batch_template':controller_template, 'c.PBSEngineSetLauncher.batch_template':engine_template })
Yes! You populate the
cluster.configobject, which is the same ascin youripcluster_config.py:cluster=ipp.Cluster( n=4, controller_ip='*', profile='pbs-2021-10-28', ) # this is the same config object you would configure in ipcluster_config.py # you don't have to call it `c`, but if you do, the rest will look familiar c = cluster.config c.PBSControllerLauncher.batch_template = controller_template c.PBSEngineSetLauncher.batch_template = engine_template await cluster.start_cluster()
Reacted by Lukas KoschmiederFantastic! Thank you!
Adding lots of examples to my documentation todo list...
I'll make an 8.0 beta tomorrow. It would be great if you could test it out!
Okay, great, I will test it.
I've got another question and potential point for the documentation todo list: How do you configure the controller dynamically in Python / Jupyter Notebook (replacement for
ipcontroller_config.py)? For instance, how would you setc.HubFactory.ip='*'?That can be
c.Cluster.controller_ipvia config, or since it's on the Cluster object, it can be a constructor argument:Cluster(controller_ip="*")
HubFactoryis removed and replaced byIPController, so if you do still have an ipcontroller_config.pyit would bec.IPController.ip = '*'`.The ambiguity is because there are really two things you are configuring:
- Cluster (and thereby Launchers) which start processes like ipcontroller, and
- ipcontroller, ipengine themselves
Some common options for configuring the controller itself can be done on the Cluster, but for the most part ipcontroller is configured directly through either
ipcontroller_config.pyorControllerLauncher.controller_args.Cluster(controller_ip="*")is really a shortcut forc.ControllerLauncher.controler_args.append("--ip=*").Reacted by Lukas Koschmieder- changed the title
[-]Unable to connect_client_sync()[/-][+]PBSControllerLauncher: Unable to connect_client_sync()[/+]on Oct 29, 2021 I just published
8.0.0b1if you could give it a tryOkay, I'm currently in the middle of something but I will give it a try this afternoon/evening.
No rush! I've only got a few more minutes of work before the weekend. I'll probably aim to do a release around the end of next week.
I've installed the new beta version
8.0.0b1and everything is looking fine - except forstart_cluster_sync, which now appears to be significantly slower. The controller starts immediately but there is some noticeable delay before the engines come up. I haven't tested if this delay scales with the number of engines. I was using 4 engines in my test.import time start = time.time() cluster.start_cluster_sync() end = time.time() print(end - start)
8.0.0b1outputJob submitted with job id: '23909' Starting 4 engines with <class 'ipyparallel.cluster.launcher.PBSEngineSetLauncher'> Job submitted with job id: '23910' 30.149987459182747.1.0(conda-forge) outputJob submitted with job id: '23911' Starting 6 engines with <class 'ipyparallel.cluster.launcher.PBSEngineSetLauncher'> Job submitted with job id: '23912' 1.15324068069458Edit: If the release is next week, unfortunately I won't be able to participate in additional beta testing because I am on holiday until Nov 8th.
except for start_cluster_sync, which now appears to be significantly slower.
That makes sense. It's the new
Cluster.send_engines_connection_envoption, which means by defaultstart_clusterwaits for the controller to finish starting before starting the engines, because the connection info is passed via environment through the Launcher. To disable this and rely on the connection files on disk (pre-8.0 behavior):cluster = Cluster(send_engines_connection_env=False, engines='pbs', controller='pbs', cnotroller_ip='*')
then the engine and controller jobs should both be submitted immediately.
@lukas-koschmieder I just published 8.0.0rc1. Can you test and then close here if you think everything is resolved?
I have been using ipyparallel 6 for a while and would like to migrate to ipyparallel 7 mainly due to fact that the new Cluster API enables you to manage the entire process through a Jupyter Notebook. Unfortunatelly, I am having difficulties to connect a client to my cluster.
I have created a new IPython profile adding a custom
ipcluster_config.py, which is a modified version of my existing/working config for ipp 6 (see below).I can successfully start a cluster spawning two PBS jobs (controller and engine).
But if I run the following line, the notebook will only show that the kernel is busy and it will never actually finish.
Am I using the API incorrectly? What might be the problem?
ipcluster_config.py