Unifi Module Removed random events

We have a Unifi site with around 43 switches. Dual ECS Campus Agg for the core, several Agg Pro’s for step-down in some areas, and many USW Pro 48/24 POE for most of the edge switches. Almost all the edge switches are connected to the core using MCLAG (all 26 groups the Campus will allow are in use, and yes, we need more).

The self-hosted UnifiOS controller reports logs to Graylog, and we have that send out Email “traps” to us, letting us know every time a port does up or down or if an SFP/DAC module is inserted or removed, and other things that the controller hides from customers.

We have noticed over the years (prior to the ECS Campus Agg and after) that we regularly and randomly have Module Removed and Module Inserted messages (along with the port going down and up). Several times a week, spread across all the switches. Thankfully, LAG and MCLAG mitigates most (but not all) of that. But the question remains- why?

It doesn’t seem to matter if they are optical SFP, DAC, AOC, what brand SFP (including FS generic, Netgear, and Unifi), time of day, day of week, speed, distance, which model switch, or what firmware versions. Some ports do seem affected more than others, but not enough to point to any particular issue with that run, compared to others. The events are always quick, usually under one second from phantom removal to insertion and port down to up. And the effects are real (the port is actually going down and then back up).

Has anyone else noticed this before? Again, if you don’t examine the Syslog logs, you might never know this is happening, because the controller doesn’t report these events, otherwise.

I have seen this behavior on our Aruba switches as well so I don’t think this is unique to unifi. We use Zabbix to monitor our switches it reports ports up/down all day long to the point we had to disable those notifications. We always chalked it up to PCs going up and down during updates, reboots etc since it mostly happened on our floor switches. Not sure if that helps just my $0.02.

1 Like

Thanks for the feedback. We are used to client ports going up/down, that is “normal” for a zillion reasons (desktop/printer/camera/etc reboot, poor ethernet patch cable or jack, power problem, etc). But it is troubling to see these with SFP switch interconnects. That is my concern- these switches are not rebooting or going up/down, just random up/down on the up/downlinks to other switches on their high-speed SFP ports.

Are you able to pull switch or even closet temps to correlate the temps at the time of outages? Those modules can get really toasty and possible the modules are not happy about it. For funzies, if you have an affected switch with an available SFP port, throw one in that’s not doing anything and see if it ever goes down. I would wager you would have more data points if the problem was at a higher level.

We have not been able to find any correlation between module temperature or room temperature with the events. All switches are affected, not just some. But some ports do seem to be affected more than others. I am planning to swap some modules in those that are the highest number of events and see if anything changes. But it is far from an exact science.