Skip to content

Planned outage for upgrade - Jun 2 during US business hours: Equinix aarch64 "Altra" systems #2948

Description

@sxa

Message from Equinix:

We would like to inform you that we will be upgrading the Mt Jade systems with additional memory (16 units x 16GB DIMM) and disk (1 unit x 1TB SSD U.2). This will ensure that the Mt Jade systems have the correct capacity as per the planned specifications (512GB RAM, 2TB SSD).

The upgrade is planned to be performed by Equinix on 2nd Jun 2022 (Thursday), in US business hours.

Activity

  1. changed the title [-]Planned outage for upgrade - Jun 2 during US business ours: Equinix aarch64 "Altra" systems[/-] [+]Planned outage for upgrade - Jun 2 during US business hours: Equinix aarch64 "Altra" systems[/+] on May 14, 2022
  2. richardlau commented on May 16, 2022

    @richardlau
    Member

    Other bits from the notice:

    Important points to be noted:

    • The upgrade of the machines will take one calendar day and during that time you will Not be able to access or use those machines.
    • Equinix Ops team should not have to remove the instances from the servers for the upgrade. But we advise you to plan backup of project data before the planned upgrade date, just in case something goes wrong and data is lost.

    Please shut the instance down, as Equinix team will need to power off the systems to add the parts.

    (Second bullet shouldn't be an issue for us as in theory the set up is all in the Ansible scripts.)

    2nd and 3rd June is a public holiday here so I'd prefer not to have to be involved. Is someone else able to handle shutting down/checking the machines come back up? cc @nodejs/build-infra

  3. richardlau commented on Jun 2, 2022

    @richardlau
    Member

    I've powered down the two Altras in preparation for tomorrow's maintenance. As mentioned I might not be around much for the next two days -- other build WG members have said they'd check the machines come back up after the maintenance is over.

  4. mhdawson commented on Jun 2, 2022

    @mhdawson
    Member

    Just waiting for the email telling us they are back up. Once that comes in I'll check if the machines are running or not

  5. mhdawson commented on Jun 2, 2022

    @mhdawson
    Member

    I had to log using the out-of-band method and then exit to get them both to come back up. The container images all seem up now so I think that should be good.

    What I'm not sure is if we also disabled some part of the CI job while they were down, will look at that a bit later today.

  6. mhdawson commented on Jun 2, 2022

    @mhdawson
    Member

    I can see there is a backlog on https://ci.nodejs.org/job/node-test-commit-arm/nodes=ubuntu2004-armv7l/ which is now being executed so don't think there is anything to re-enable.

  7. mhdawson commented on Jun 2, 2022

    @mhdawson
    Member

    https://ci.nodejs.org/computer/test-equinix-ubuntu2004-arm64-1/ is still offline but from the load stats I don't think its been used recently -https://ci.nodejs.org/computer/test-equinix-ubuntu2004-arm64-1/load-statistics?type=hour

    Instead it is likely just the underling host that we have the containers running on.

  8. sxa commented on Jun 3, 2022

    @sxa
    MemberAuthor

    @mhdawson DId you have to kick it from the out of band console again or was just powering them back in the web UI adequate?

  9. mhdawson commented on Jun 13, 2022

    @mhdawson
    Member

    @sxa I had to kick it from the out of band console.

  10. richardlau commented on Jun 13, 2022

    @richardlau
    Member

    Closing as the upgrade happened. #2894 is still ongoing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions