We have a Unifi site with around 43 switches. Dual ECS Campus Agg for the core, several Agg Pro’s for step-down in some areas, and many USW Pro 48/24 POE for most of the edge switches. Almost all the edge switches are connected to the core using MCLAG (all 26 groups the Campus will allow are in use, and yes, we need more).
The self-hosted UnifiOS controller reports logs to Graylog, and we have that send out Email “traps” to us, letting us know every time a port does up or down or if an SFP/DAC module is inserted or removed, and other things that the controller hides from customers.
We have noticed over the years (prior to the ECS Campus Agg and after) that we regularly and randomly have Module Removed and Module Inserted messages (along with the port going down and up). Several times a week, spread across all the switches. Thankfully, LAG and MCLAG mitigates most (but not all) of that. But the question remains- why?
It doesn’t seem to matter if they are optical SFP, DAC, AOC, what brand SFP (including FS generic, Netgear, and Unifi), time of day, day of week, speed, distance, which model switch, or what firmware versions. Some ports do seem affected more than others, but not enough to point to any particular issue with that run, compared to others. The events are always quick, usually under one second from phantom removal to insertion and port down to up. And the effects are real (the port is actually going down and then back up).
Has anyone else noticed this before? Again, if you don’t examine the Syslog logs, you might never know this is happening, because the controller doesn’t report these events, otherwise.