New on LowEndTalk? Please Register and read our Community Rules.
All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.
All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.
Comments
hmmmmm

Isn't onidel HA?
The HA cluster went down.
Ya I am saying it more of a general sense and NOT targetting Onidel.
Even with HA support from Onidel the whole location could be down for instance, so there's need for cross-region DR etc but it just isnt easy to have a backup for everything.
HA isn't a DR;
Raid isn't a backup;
Snapshot isn't a backup;
Better make backups
Do daily offsite backups of important stuff to Hetzner. And I know it's morning here, I'm here too... Alarm set for 6AM M-F
Was wondering why I woke up with no notifications on my phone... My personal VPN is hosted on the same network that was down.. Womp womp. No internet connectivity to phone, no notifications to address when I wake up haha.
I didn't want to go to the gym this morning anyway.
Seems that now most of my Singapore VMs are migrated and stable, at least the most important ones.
So, now that I can do work today, I guess I'll head to the office and figure out how to make something like this occurring again impact me less. Main thing that would have saved me is some form of BGP failover to another regional BGP enabled VM. Probably won't fuss with much beyond making sure I can maintain a wireguard connection in a failover state, the rest can be DNS based failovers for now.
is this what we call clusterf**k?
All SGP VMs are now back online. If you are still experiencing any issues, please open a support ticket and we will take a look.
Betterstack paged us at 5:16 AM AEDT about the issue. Nothing better than an critical incident at this time on a public holiday.
The bad is, of course, the downtime was much longer than we expected. This incident took down the entire cluster, and we ran into several issues during the recovery that slowed the process down. It was frustrating every minute that @oloke and I were working to bring the cluster back up, because we knew we had failed to deliver the service reliablity and availability we expect from ourselves. There is no excuse for that - this was simply not up to our standard.
The good news is that there was no data loss. We have also already identified several improvements that will definitely come out of the PIR, although there are still some issues we need to investigate further. We will share more details with impacted customers once we have a clearer understanding of what happened and what we are doing to prevent it from happening again.
FWIW, the control panel were taken offline intentionally to prevent customer actions from interfering with the recovery process. They were brought back online once we introduced the In Migration status.
I'd be interested to read a postmortem.
The things that I remember a few second before SSH being disconnected is Load average on htop becaming 7 to 9 on my 2 vcpu Turin vps.
Asking @onidel via ticket and answered nothing wrong then next message is there was down.
For me, this the first time experiencing with the issue after more than 1 year using Onidel, since I have Turin and Milan vps, also block storage and object storage, all Singapore.
I think my Turin vps migrated to Milan temporarely
lscpu
Architecture: x86_64
CPU op-mode(s): 32-bit, 64-bit
Address sizes: 48 bits physical, 48 bits virtual
Byte Order: Little Endian
CPU(s): 2
On-line CPU(s) list: 0,1
Vendor ID: AuthenticAMD
Model name: AMD EPYC 7543 32-Core Processor
Same here. I've been thinking about setting up a cluster for a long time, but I always doubted something like this would happen, having complex setup has their risk as well, so I want to know exactly what happened as well. I admire OniDel for having such a nice setup
Sorry guys I started using all my idlers with Onidel in SG that caused all nodes to be overloaded
My guesses:
You forgot
DNS
Ah right, but I'd assume it's statically configured
Only the weaks have backup, strong men redo the shit when it lost.
I was tinkering with UKI and Measured boot just moment before this happened...but I did it in Sydney.
My compute is up for renewal, and I definitely won't fall for the 'high availability' pitch again.
Yeah you should pay thousands to big corp instead, that will be "always availability" for you.
I’ll just point out that any provider can have major issues. Any provider.
AWS has had major issues. GCP has also had major issues.
I have confidence that Onidel and their team have learned a lot and the PIR will be very interesting to read, be educational and will take steps to improve and grow.
What will be the test is what they do after that and what they implement. Which again I’m confident that it will be preventative and will improve their already pretty good service (not a customer right now, was but I had 0 issues when I was).
I would trust them.
Yes. A provider you can touch, not "talk to support agent"
Yeah let me just throw this rock over the ditch. I’m sure it will make it!
Sorry, your rock dissolved in the acid
Can you transfer it to me free or charge?
Aw man…
If it is one of those promo I will be happy to take it from you
Yeah unfortunately there is no availability when the entire cluster is down
We are close to completing our investigation and are confident that we now understand what happened. We are currently reviewing the actions taken during the incident to identify anything we could have done differently to restore service more quickly.
We will share the PIR once it is ready.
Juz wan to add sugar coat im a s3 onidel user. Pointed some ui pagination n breadcrumb bug in ticket, bang next i know the app team working on fixing the bugs. Within 24 hrs i think, onidel ask me to reconfirm if ui bugs fixed. Thumb up for honest genuine provider.
Hey guys, lurker here, finally jumping in.
Saw SG just had a restock / opened up some slots. Anyone know what the expansion roadmap looks like for SG capacity? Tempted to grab a couple of boxes, but I'm worried about hitting a wall when it's time to scale up once the current node fills up.
Waiting for HCMC stock 🇻🇳