Howdy, Stranger!

It looks like you're new here. If you want to get involved, click one of these buttons!


New on LowEndTalk? Please Register and read our Community Rules.

All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.

Onidel Cloud's Thread - Announcements, Feedbacks and Discussions!

1…151617181921»

Comments

  • hmmmmm

  • @FAT32 said:
    HA is easy until you actually need to do it

    Isn't onidel HA?

  • @Motion3549 said:

    @FAT32 said:
    HA is easy until you actually need to do it

    Isn't onidel HA?

    The HA cluster went down. >:)

  • FAT32FAT32 Administrator, Deal Compiler Extraordinaire

    @Motion3549 said:

    @FAT32 said:
    HA is easy until you actually need to do it

    Isn't onidel HA?

    Ya I am saying it more of a general sense and NOT targetting Onidel.

    Even with HA support from Onidel the whole location could be down for instance, so there's need for cross-region DR etc but it just isnt easy to have a backup for everything.

    Thanked by 1Motion3549
  • rpqurpqu Member


    HA isn't a DR;
    Raid isn't a backup;
    Snapshot isn't a backup;
    Better make backups

  • MannDudeMannDude Patron Provider, Veteran

    @rpqu said:

    @MannDude said:

    @rpqu said:

    @MannDude said:

    @rpqu said:

    @MannDude said:
    I still don't understand why people host their company website on their own infrastructure, or put it up without a simple static failover solution to a "Ah sheiiitttt, we'll be right back" type placeholder message of sorts.

    Because they're confident with their own infras?

    Well, seems to be going great now, right? A cheap Digital Ocean or similar VPS is like $5/mo and is the difference between a customer being able to access the company's main website to see if something is wrong or hunt LowEndTalk or other public forums to see if there is a thread with a link to a status page somewhere. You can still be confident while planning for such events.

    But, it doesn't matter. I found this thread. Others, may not find it and they'll make whatever assumptions people may make when things are down and the company website is down. $5/mo for a failover VPS could be the difference between a customer's angry LET/LES/WHT/X or other post about not knowing whats going on and a calm, informed customer.

    Well, that's true. And more reasons for LET provider to collect vps from other providers.

    Also more reason for me to do the same too, haha.

    Now I need second SGP provider with BGP at the minimum so I can do BGP failover to them or something.... I'll have to look more into that, something new to me.

    So, while I preach about all eggs in one basket I'm guilty of doing the same. I have a SEA startup that relies on Onidel infra but besides some sites being served direct from Bunny, some of the backend/platform stuff is down still due to the announcement still being down. But that's an oversight and failure on my part to not plan accordingly.

    LOL LMAO. At least tell me you have WAL or hourly offsite backup. It's morning on SEA tz. So, better deploy it somewhere
    As for BGP hmmm.

    Do daily offsite backups of important stuff to Hetzner. And I know it's morning here, I'm here too... Alarm set for 6AM M-F :)

    Was wondering why I woke up with no notifications on my phone... My personal VPN is hosted on the same network that was down.. Womp womp. No internet connectivity to phone, no notifications to address when I wake up haha.

    I didn't want to go to the gym this morning anyway.

    Seems that now most of my Singapore VMs are migrated and stable, at least the most important ones.

    So, now that I can do work today, I guess I'll head to the office and figure out how to make something like this occurring again impact me less. Main thing that would have saved me is some form of BGP failover to another regional BGP enabled VM. Probably won't fuss with much beyond making sure I can maintain a wireguard connection in a failover state, the rest can be DNS based failovers for now.

    Thanked by 1rpqu
  • budi1413budi1413 Veteran

    is this what we call clusterf**k? :D

    Thanked by 2Obelous rpqu
  • onidelonidel Member, Patron Provider, Top Host, Megathread Squad

    All SGP VMs are now back online. If you are still experiencing any issues, please open a support ticket and we will take a look.

    Betterstack paged us at 5:16 AM AEDT about the issue. Nothing better than an critical incident at this time on a public holiday.

    The bad is, of course, the downtime was much longer than we expected. This incident took down the entire cluster, and we ran into several issues during the recovery that slowed the process down. It was frustrating every minute that @oloke and I were working to bring the cluster back up, because we knew we had failed to deliver the service reliablity and availability we expect from ourselves. There is no excuse for that - this was simply not up to our standard.

    The good news is that there was no data loss. We have also already identified several improvements that will definitely come out of the PIR, although there are still some issues we need to investigate further. We will share more details with impacted customers once we have a clearer understanding of what happened and what we are doing to prevent it from happening again.

    FWIW, the control panel were taken offline intentionally to prevent customer actions from interfering with the recovery process. They were brought back online once we introduced the In Migration status.

  • forestforest Member

    I'd be interested to read a postmortem.

  • kidoskidos Member

    The things that I remember a few second before SSH being disconnected is Load average on htop becaming 7 to 9 on my 2 vcpu Turin vps.

    Asking @onidel via ticket and answered nothing wrong then next message is there was down.

    For me, this the first time experiencing with the issue after more than 1 year using Onidel, since I have Turin and Milan vps, also block storage and object storage, all Singapore.

    I think my Turin vps migrated to Milan temporarely
    lscpu
    Architecture: x86_64
    CPU op-mode(s): 32-bit, 64-bit
    Address sizes: 48 bits physical, 48 bits virtual
    Byte Order: Little Endian
    CPU(s): 2
    On-line CPU(s) list: 0,1
    Vendor ID: AuthenticAMD
    Model name: AMD EPYC 7543 32-Core Processor

  • HostDZireHostDZire Patron Provider, Veteran

    @forest said:
    I'd be interested to read a postmortem.

    Same here. I've been thinking about setting up a cluster for a long time, but I always doubted something like this would happen, having complex setup has their risk as well, so I want to know exactly what happened as well. I admire OniDel for having such a nice setup :)

  • FAT32FAT32 Administrator, Deal Compiler Extraordinaire

    Sorry guys I started using all my idlers with Onidel in SG that caused all nodes to be overloaded

  • rpqurpqu Member

    @forest said:
    I'd be interested to read a postmortem.

    My guesses:

    • split-brain
    • network problems
    • clock drifts
    Thanked by 2FAT32 nullnothere
  • FAT32FAT32 Administrator, Deal Compiler Extraordinaire

    @rpqu said:

    @forest said:
    I'd be interested to read a postmortem.

    My guesses:

    • split-brain
    • network problems
    • clock drifts

    You forgot

    DNS

  • rpqurpqu Member

    @FAT32 said:

    @rpqu said:

    @forest said:
    I'd be interested to read a postmortem.

    My guesses:

    • split-brain
    • network problems
    • clock drifts

    You forgot

    DNS

    Ah right, but I'd assume it's statically configured :D

    Thanked by 1FAT32
  • nghialelenghialele Member

    @rpqu said: Better make backups

    Only the weaks have backup, strong men redo the shit when it lost.

  • I was tinkering with UKI and Measured boot just moment before this happened...but I did it in Sydney.

Sign In or Register to comment.