Wednesday, July 22, 2026
[Incident Report #037][vIX] Platform Issues and Tooling….
What are Incident Reports?
As a community‑operated and governed virtual internet exchange, FurrIX maintains
a commitment to open and honest communication with its members. Sometimes
during the operation of the exchange, we run into issues that impact the state of
the exchange and member connectivity. When this happens, the FurrIX operations
team publishes an incident report to ensure all members remain informed. As a
hobbyist‑rooted vIX, we aim to keep communication clear, accessible and practical
to the best of our ability.
Summary
On July 22nd, 2026 at approximately 12:45 EST, the FurrIX vIX entered a degraded
operational state. We lost access to PHYONE’s Proxmox control plane and experienced
partial failure of NS1’s DoT/DoH services. The root cause was a failure in newly
developed SSL certificate distribution tooling, which corrupted certificate stores on
both PHYONE and NS1. This resulted in service outages and exposed gaps in our
recovery lifelines.
What Happened?
Our lead engineer was developing new automation to handle SSL certificate
distribution across the vIX, aiming to reduce manual volunteer workload.
During a trial run on the distro network, the tooling behaved unexpectedly
and corrupted the certificate stores on both PHYONE and NS1.
Furthermore, when it was noticed that we no longer had UI access, our volunteers
tried using the recovery options we had thought were fully in place and found out
very quickly that neither lifeline was fully operational. For roughly two hours, the
vIX was running headless with limited administrative control.
We were seeing the following issues:
- Control Plane Damage — Proxmox UI on PHYONE offline
- NS Partial Failure — NS1’s DoT/DoH endpoints failing
- RNDC and other control‑plane operations to become unstable
The code appeared correct during review, but a permissions issue slipped through and
only manifested once deployed. A right fucken oops, that was. Compounding the issue,
when volunteers attempted to use our recovery lifelines, we discovered that neither of
them were fully operational. For roughly two hours, the vIX was running headless with
limited administrative control.
What did we do to fix this?
We learned that while our OOB recovery network was operational, our recovery shell
accounts had not been setup on PHYONE. We promptly reached out to the data center
for a KVM to be put on the machine as soon as possible. Upon getting that setup, our
volunteers deployed our recovery account via the machine shell and then detached from
KVM to continue the recovery as to our DRP for ‘Access Loss, PHYONE’ and we modified
our SSL handling scripts to ensure that NS1 set the correct permissions on the cert store
and that PHYONE now validates certs before installing them and reloading services.
During the recovery, we also learned the same issue was affecting tooling on NS1 itself and
we got that sorted to make sure it also validates its certs and keys before reloading any
services.
While the recovery was ongoing, we had a few blips in networking as the KVM was connected
and disconnected from PHYONE- but we should be fully alive and working again. Oh yea, we also
added Discord alerts to our scripts so that we are able to keep an eye on their execution and
catch any problems that might occur.
Response & Recovery Actions
1. Restoring Access to PHYONE
- Verified OOB recovery network was operational
- Discovered recovery shell accounts were not configured
- Contacted the data center and requested KVM attachment
- Once KVM was online, deployed recovery account via local shell
- Detached KVM to minimize network blips and continued
recovery per DRP “Access Loss, PHYONE”
2. Repairing Certificate Store PHYONE
- Identified corrupted cert stores on PHYONE
- Updated SSL handling scripts to:
— Validate certs before installation
— Validate key/cert pairing
— Reload services only after validation passes
3. Fixing NS1 Tooling and Cert Store
- Found identical validation/permission issues affecting NS1’s tooling
- Updated scripts to ensure NS1 validates certs and keys before reloading services
- Apply correct permissions (bind:bind)
Root Cause
A permissions and validation oversight in new SSL distribution tooling caused certificate
corruption on PHYONE and NS1. Lack of fully configured recovery accounts delayed restoration.
Lessons Learned
Going forward with the development and upkeep of the FurrIX vIX,
our volunteers will be applying the following lessons to future scripts
and any custom tooling:
- Automation touching cert stores must validate before overwrite
- Permissions must be explicitly set every time
- Recovery accounts must be deployed and tested on all hypervisors
- Monitoring hooks (Discord alerts) are essential for early detection
For onlookers wondering why this was not sandboxed more aggressively before
being put into production, it is a hard truth that FurrIX does not have a replicated
environment offsite for testing these kinds of tools. Meaning a lot of the time, we
are heavily crawling our own tooling before we deploy and sometimes not so obvious
issues can crop up that our volunteers haven’t thought about before hand.
Wednesday, July 1, 2026
[Transparency Report #012][Documentation] vIX Monitoring
What are Transparency Reports?
As a community‑operated and governed virtual internet exchange, FurrIX maintains
a commitment to open and honest communication with its members. From time to
time, operational work may occur that affects the exchange or its supporting infrastructure.
When this happens, the FurrIX operations team publishes a transparency report to
ensure all members remain informed. As a hobbyist‑rooted vIX, we aim to keep
communication clear, accessible and practical to the best of our ability.
What Happened
During the MFN to FurrIX migration, a number of larger infrastructure projects took priority
and were using our volunteer’s free time to get the exchange ready for full operation.
As a result, the monitoring stack (LibreNMS + graph export scripts) fell out of sync with the
new network layout. A stale firewall rule on PHY Two’s edge router blocked the monitoring
server’s requests with changes to new PI space, causing all transit graphs to stop updating.
Because this was a volunteer‑run transition with limited available time, the issue persisted
longer than usual, roughly four months, while other critical work was completed.
Changes to the exchange:
The outdated firewall rule was corrected, restoring connectivity between the web server
and LibreNMS. Once access was restored, all graph‑generation scripts came back online
and were patched with new tooling bits for extended monitoring internally and public
facing. All vIX flow‑rate graphs are now current and visible again.
Are exchange operations affected?
Both volunteers and members now have full visibility into how the vIX carries data and
how usage trends evolve over time. Aside from improved monitoring, normal operations
continue as expected.
Saturday, June 20, 2026
[Transparency Report #010][OPERATIONS] BGP gets a Tune Up!
What are Transparency Reports?
As a community‑operated and governed virtual internet exchange, FurrIX maintains
a commitment to open and honest communication with its members. From time to
time, operational work may occur that affects the exchange or its supporting infrastructure.
When this happens, the FurrIX operations team publishes a transparency report to
ensure all members remain informed. As a hobbyist‑rooted vIX, we aim to keep
communication clear, accessible and practical to the best of our ability.
What is happening?
This week’s changes focused on tightening routing policy between Edge, our
member access routers and the services router (Catos, Ikus and Nardoragon).
The goal was simple:
- eliminate any possibility of route leaks
- enforce strict prefix‑origination rules
- ensure the exchange remains hobbyist‑grade, stable, and predictable
All required changes were applied without service interruption. All member routes
remained visible and stable throughout the transition.
Changes to the exchange:
Our volunteers have implemented uniform BGP filtering across all internal routers.
Catos, Ikus, and Nardoragon:
- May only advertise their assigned /58 prefixes
- May only learn the default route from Edge
- Cannot advertise our PI /45 or /46 aggregate anywhere
- Cannot learn leaked routes from Edge or from each other
Edge:
- Only advertises ::/0 toward all downstream routers
- Only accepts each downstream router’s assigned /58
- Is the only router permitted to originate the /45 and /46 aggregates
- Will only originate those aggregates once we obtain our own ASN (maps already in place)
Prefix‑lists and route‑maps have been standardized across all routers to ensure the fabric
remains predictable and safe for our volunteers and members to continue learning and
experimenting within the exchange. This includes consistent permit/deny ordering, strict
prefix matching and hardened default‑deny behavior.
Are exchange operations affected?
Everything is operating as expected. This was much‑needed work in the background to ensure
long‑term stability and predictability of the exchange. These changes make our BGP setup more
oops‑proof, better hardened and more aligned with real IX operational practices — while still
keeping the environment friendly for hobbyist experimentation.
Tuesday, June 16, 2026
[Transparency Report #009][OPERATIONS] BGP Is Enabled! (Internally)
What are Transparency Reports?
As a community‑operated and governed virtual internet exchange, FurrIX maintains
a commitment to open and honest communication with its members. From time to
time, operational work may occur that affects the exchange or its supporting infrastructure.
When this happens, the FurrIX operations team publishes a transparency report to
ensure all members remain informed. As a hobbyist‑rooted vIX, we aim to keep
communication clear, accessible and practical to the best of our ability.
What is happening?
This is a good thing for the exchange to have figured out. As of Jun 15th, we have learned
how to configure and enable BGP on OpnSense within the exchange. This means our techs
can now peer the exchange with member delegated /64s over /127 wireguard links! This is
a goal that we have been working towards, which also serves to get us moving towards our
goal of one day having a public ASN. Going forward, members who join the exchange will
have the option of having their /64 on-link or BGP peering with us an announcing their
/64 to our routing fabric.
Changes to the exchange:
- FurrIX Transit Fabric: Edge, Catos and Nardoragon are all peered using AS65300. Edge
announces a default route downstream, while the other two routers announce their assigned
/58s to the Edge.
- Exchange Member Peering: FurrIX has reserved AS65320 for peering with members of
our exchange, we also have started to rework our peering policies along with reserving
AS65400-65500 for member BGP sessions and AS65501-AS6550 for peering with other
hobbyist networks.
Changes Proposed:
Eventually FurrIX would like to add a BGP looking glass to our network that is peered
with the Edge that will should all ASNs and routes on the exchange, but this is a ways
off for the moment.
Are exchange operations affected?
Everything is operating normally, this was just quiet work in the background in order to
mature the exchange a little further and get to a point that we are reaching some of our
goals that were set for this year.
Monday, May 25, 2026
[Transparency Report #007][OPERATIONS] Full Environment Rebuild Scheduled WIP
What are Transparency Reports?
As a community‑operated and governed virtual internet exchange, FurrIX maintains
a commitment to open and honest communication with its members. From time to
time, operational work may occur that affects the exchange or its supporting infrastructure.
When this happens, the FurrIX operations team publishes a transparency report to
ensure all members remain informed. As a hobbyist‑rooted vIX, we aim to keep
communication clear, accessible and practical to the best of our ability.
What is happening?
The FurrIX vIX is currently going through its rebuild of our exchange and it is taking a little
longer than we expected. Due to a miscommunication, reinstalling the physical server’s OS
took a bit of time.
What has been reworked so far:
- Phy One: The ProxMox host has been rebuilt
- Core Router: We condensed our IPv6 edge and core router into one VM
- Nardoragon Router: Our services router is back online with new config
- Catos vIX Access Router: Has been pulled from backup and reconfigured
- NS1/Games-3P: These member facing services are back online
- Web Server: Our websites are back online
Parts of the exchange still being worked on:
- Mail-NG: the mail server has to be brought back online
- Ikus vIX Access Router: Secondary member facing router still being reconfig’d
- NMS: We currently have no monitoring, needs to be reconfigured
Are exchange operations affected?
Yes — temporarily.
During the rebuild window, routing and service availability will be null as systems are rebuilt
and renumbered. Once the work is complete, normal operations will resume with improved
stability, ease of expansion, better rooted upkeep and clarity.