New on LowEndTalk? Please Register and read our Community Rules.
All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.
All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.
Comments
hmmmmm

Isn't onidel HA?
The HA cluster went down.
Ya I am saying it more of a general sense and NOT targetting Onidel.
Even with HA support from Onidel the whole location could be down for instance, so there's need for cross-region DR etc but it just isnt easy to have a backup for everything.
HA isn't a DR;
Raid isn't a backup;
Snapshot isn't a backup;
Better make backups
Do daily offsite backups of important stuff to Hetzner. And I know it's morning here, I'm here too... Alarm set for 6AM M-F
Was wondering why I woke up with no notifications on my phone... My personal VPN is hosted on the same network that was down.. Womp womp. No internet connectivity to phone, no notifications to address when I wake up haha.
I didn't want to go to the gym this morning anyway.
Seems that now most of my Singapore VMs are migrated and stable, at least the most important ones.
So, now that I can do work today, I guess I'll head to the office and figure out how to make something like this occurring again impact me less. Main thing that would have saved me is some form of BGP failover to another regional BGP enabled VM. Probably won't fuss with much beyond making sure I can maintain a wireguard connection in a failover state, the rest can be DNS based failovers for now.
is this what we call clusterf**k?
All SGP VMs are now back online. If you are still experiencing any issues, please open a support ticket and we will take a look.
Betterstack paged us at 5:16 AM AEDT about the issue. Nothing better than an critical incident at this time on a public holiday.
The bad is, of course, the downtime was much longer than we expected. This incident took down the entire cluster, and we ran into several issues during the recovery that slowed the process down. It was frustrating every minute that @oloke and I were working to bring the cluster back up, because we knew we had failed to deliver the service reliablity and availability we expect from ourselves. There is no excuse for that - this was simply not up to our standard.
The good news is that there was no data loss. We have also already identified several improvements that will definitely come out of the PIR, although there are still some issues we need to investigate further. We will share more details with impacted customers once we have a clearer understanding of what happened and what we are doing to prevent it from happening again.
FWIW, the control panel were taken offline intentionally to prevent customer actions from interfering with the recovery process. They were brought back online once we introduced the In Migration status.
I'd be interested to read a postmortem.
The things that I remember a few second before SSH being disconnected is Load average on htop becaming 7 to 9 on my 2 vcpu Turin vps.
Asking @onidel via ticket and answered nothing wrong then next message is there was down.
For me, this the first time experiencing with the issue after more than 1 year using Onidel, since I have Turin and Milan vps, also block storage and object storage, all Singapore.
I think my Turin vps migrated to Milan temporarely
lscpu
Architecture: x86_64
CPU op-mode(s): 32-bit, 64-bit
Address sizes: 48 bits physical, 48 bits virtual
Byte Order: Little Endian
CPU(s): 2
On-line CPU(s) list: 0,1
Vendor ID: AuthenticAMD
Model name: AMD EPYC 7543 32-Core Processor
Same here. I've been thinking about setting up a cluster for a long time, but I always doubted something like this would happen, having complex setup has their risk as well, so I want to know exactly what happened as well. I admire OniDel for having such a nice setup
Sorry guys I started using all my idlers with Onidel in SG that caused all nodes to be overloaded
My guesses:
You forgot
DNS
Ah right, but I'd assume it's statically configured
Only the weaks have backup, strong men redo the shit when it lost.
I was tinkering with UKI and Measured boot just moment before this happened...but I did it in Sydney.