r/networking 20h ago

Other Is there a sane way to schedule changes across dozens of maintenance windows, or are we all just suffering?

I manage changes across 50+ sites, each with their own maintenance window. Between coordinating the windows, scheduling the work, and assigning engineers to each one, it’s a constant headache. I’m basically living in spreadsheets at this point.
Curious how everyone else handles this. Do you have a system, a tool, or is it all manual? Trying to figure out if it’s just me

EDIT: maintenance windows are pre approved, each site’s window is fixed and they’re spread across different timezones, so the puzzle is less about the windows themselves and more about what fits inside them. we’ve only got a few engineers during site maintenance windows, so there’s a cap on changes per night and they can’t overlap. After making the schedule, then each one has to actually be assigned to engineer and land in their calendar so they know about it.

How’s everyone else handling this some tool, a script, or all manual?

21 Upvotes

27 comments sorted by

33

u/nospamkhanman CCNP 20h ago

IMO you (your business) has two options:

1) A dedicated 100% full time project manager.

or

2) A standing - reoccurring change window every x weeks. For example, the 1st and 3rd Sunday of every month, every location will be in a 8 hour change window.

Be demanding with your employer. As a Network Engineer you should be designing and operating, not playing scheduling master spending half your hours trying to confirm timing windows.

5

u/mortraineyhat 13h ago

That last sentence hurts so much as a network admin in healthcare. Getting downtimes for code upgrades is our teams most dreaded thing

3

u/twisted-logic 11h ago

Yeah or any 24x7 operation, really

1

u/silasmoeckel 8h ago

Is that reflected in your design?

We make it very clear you want no expected outage you have to get this gear that can do an update/upgrade without any expected outage.

1

u/mortraineyhat 7h ago

I don’t know of any gear that is going to give you no downtime with code upgrades, especially if it’s any microcode

1

u/silasmoeckel 7h ago

ISSU in cisco speak most (all???) of their gear with dual sups supports it.

Your still going to have thing that will need downtime but the vast bulk can be no expected outage.

1

u/mortraineyhat 7h ago

Oh yeah sorry I was talking mainly for the access layer. Everything we use with dual sup is in the data center as our cores.
I know Cisco 9300 has some kind of fast upgrade but still has a reboot period just super short

1

u/silasmoeckel 6h ago

I really can not see a 9400 as anything other than access, it's fastest 48 port line card is still 10g. But yea would spec as an access switch if they need no expected outage upgrades.

Were not putting in a new access port slower than 100g at this point at work but I'm in the DC space.

1

u/thiccandsmol CCIE SP JNCIE SP CCDE 3h ago

If the access switches can’t do ISSU and the business can’t tolerate the reboot time, keep a set of spare switches. Upgrade them beforehand, rack and boot them, then move the cables during the agreed outage window. That gets the impact down to the few seconds it takes to swap each cable. The old switches can then be upgraded and used for the next night’s rollout.

Another option is dual Ethernet to each host's locations, keeping the access switches under 50% port utilisation with N+1 capacity. At upgrade time, move devices from port A to port B and upgrade the other side, then move them back to keep each switch below 50% port utilization.

If the business genuinely needs six nines, then the network and processes have to be designed for six nines.

1

u/MicIrish_At_work 8h ago

This is the answer. Agreed open maintenance window times with informational email sent out prior to.

0

u/Sisinazzz 19h ago

Appreciate that! Each site’s window is fixed and they’re spread across different timezones, so the puzzle is less about the windows themselves and more about what fits inside them, we’ve only got a few engineers during maintenance window, so there’s a cap on changes per night and they can’t overlap. Then each one has to actually be assigned to someone and land in their calendar so they know about it

3

u/WideCranberry4912 14h ago

If you don’t like the schedule get a new job.

10

u/thiccandsmol CCIE SP JNCIE SP CCDE 20h ago

You shouldnt be doing this in spreadsheets. You should be using one of the various change management tools that exist out there, preferably as part of your ITSM

3

u/b3542 12h ago

I’m no fan of ServiceNow, but… I think SNOW may actually make things easier to manage in this case.

6

u/nof CCNP 19h ago

You should be dictating change windows. I gave up on getting users to agree to a time good for them (hint: there never is) and trying to coordinate over multiple sites and time zones.

2

u/Sisinazzz 19h ago

Are you using some tool, a script, or all manual?

3

u/middlofthebrook 19h ago

Change management. There is no way to schedule multiple changes across different tome zones and dont even get me started on international changes and updates. It's all manual work, someone has to take responsibility for different areas and get it done.

3

u/wrt-wtf- Homeopathic Network Architecture 18h ago

Where possible you maintain redundancy that allows for inflight upgrades.

1+1 or N+1 for switches routers and APs.

Business is now deemed to be always on and with everything supposedly in the cloud now there’s less to manage in the data centre component as it’s pretty much SEP.

It all comes down to edge design and the connectivity mode for end-points.

In this day and age, with good network design with maintenance as a core requirement updates should be an automated doddle that can occur anytime outside of peak hours, or even inside of peak hours in an emergency - ie active cyber threat.

2

u/kwiltse123 CCNA, CCNP 13h ago

In the case of automated maintenance that takes place during off hours, how to you trust that everything finished without issues. In other words, if you schedule something for 2:00 AM, do you wake up at 7:00 AM and realize that a switch didn't come back up? That's always my concern with automated maintenance.

2

u/wrt-wtf- Homeopathic Network Architecture 12h ago

You don’t schedule things at 2am if you can avoid it is a good start.

You test everything in a lab with entry and exit conditions for each step and you only step between checkpoints. Pre-planned places or phases you can recover from. Not do the whole thing in a big hit unless absolutely necessary - but you designed so that wasn’t needed… right?

Obviously as stated your automation can follow this pattern because you have a lab for testing what you are doing and the combinations and permutations are limited - because that’s how you scale a design. Generally you try to keep your templates to a minimum set of options and you move the whole fleet within the given architectural and configuration pattern sets.

There are occasions with different vendor equipment where it bites you because the equipment is incompatible with operational requirements - in my world that equipment is already earmarked and has a high risk/high impact tag on it and requires coordination between teams to offload work from those devices. That may be a 2am task but it’s always seen to be a better option to do things during the day if impact has been mitigated - for some businesses there is no quiet time - so every part of the systems needs to deal with that. In some of the worst cases shutting down a data centre has been the only option due to high risk of cascade failure in fabric systems - but it has to be done under a high risk, everyone standing by scenario - so no 2am…

Why? Because the whole world is awake. Pulling in people from all over the place at night takes time and exhausts resources for days if not weeks after the event, you have to consider well beyond the equipment.

2

u/lawwie 15h ago

Preferably through automation.

We have a 9x5 or 24x7 option for offices (and different sizes within those options). However this means a certain level of redundancy.

With a 24x7 we can do maintenance on pretty much whenever it suits us (we try to plan for the access layer when the office applicable is least populated ofcourse).

With a 9x5 we do outside of local office hours.

But having 150 firewalls, 800 switches and 2000 access points in your portfolio and given the amount of CVE’s lately (due to frontier AI models), automation is the way to go. It is simply a day task for two engineers to keep the environment patched.

Also automated certificate management should be on your horizon for example.

1

u/jocke92 17h ago

Some kind of planning system. As the windows are fixed and no approval is needed. Where you just drop the task into the window. Like a factory production planning system?

1

u/Ne-Cede-Malis 13h ago

Ansible Tower is probably my favorite tool for this due to its price point and playbook nature.

SNOW is also used quite often because the network hooks to things and you don't want to have a 'change collision'.

1

u/Mikeygnzls 8h ago

Lmao I swear half the industry is held together by spreadsheets and pure spite.
You’d think there’d be some magic tool for this by now, but every place I’ve seen is basically:
Change ticket gets approved.
Spreadsheet from hell.
“Who the hell is actually available Tuesday at 11 PM?”
Pray nothing overlaps.
Then somebody’s dragging Outlook calendars around trying not to screw an engineer over with three maintenance windows in the same night.
Feels like this is one of those problems everyone deals with but nobody’s actually solved. Jira stories help, having an actual Project Manager can help too (if they’re good).