New on LowEndTalk? Please Register and read our Community Rules.
All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.
All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.
Comments
So here's what I think is happening.
SSDs can't overwrite flash data in place. The flash is written in relatively small pages, roughly 16 KB, but it has to be erased in much larger blocks, several MB at a time. So when a file changes, the drive writes the new data somewhere else and leaves the old data there until it can be cleaned up. The FTL (flash translation layer) tracks where the current data actually is.
Eventually the drive runs out of clean space and has to do garbage collection. It takes a block, copies out anything that's still valid, then erases the block so it can be used again. If most of the data in that block is already invalid, that's not a big deal. The problem is that this drive has had 887 TB written to it, mostly from lots of small, random writes from multiple VMs at the same time. That tends to scatter valid data, so the drive has to move more data around before it can free up space.
That also fits with what I'm seeing in the diagnostics. Reads are still fast at
8.5 GB/s, so this isn't a general performance problem. Writes are down at185 MiB/s, and individual writes were taking around 147 ms at the drive level. The identical drive next to it was doing the same work in about 10 ms.The other checks look normal. SMART is clean, wear is only 2%, temperatures are fine, the PCIe link is running at full Gen4 x4, and the power state is normal. The drive also has the same model and firmware as the one that's performing normally.
I'm going with a secure format (
nvme format --ses=1) as a potential fix. That should clear the drive's internal mapping and erase the blocks, effectively giving the drive a clean start.Since I can't look inside the FTL and prove garbage collection is causing the slowdown, this remains a diagnosis rather than something I can directly confirm. But given the symptoms and everything else I've checked, it's the most plausible explanation.
If the drive is still slow after a full secure wipe, then I'll treat it as a hardware or firmware problem. I'll put in a new drive and send the old one back under warranty.
I'll be sending out a maintenance notice tonight and get this taken care of ASAP.
No need to do that. It's definitely the fault of write amplification as you surmised, but there's already TRIM which indicates which blocks are free to the garbage collector. My guess is your configuration is such that VMs are not passing through TRIM commands, so all the free space on everyone's VM is being treated as in-use and those blocks are unavailable for allocation. Without taking advantage of TRIM, you'll see this performance degradation.
Start up one of your OS templates on that node and run
fstrim -avas root and, if its output is silent or it doesn't return a message like/: 3.1 GiB (3370385408 bytes) trimmed on /dev/vda1, then you've found your problem.Even if a product is at deep discount, if it doesn't meet what it was advertised as that constitutes false advertising. Beyond the illegality of misrepresenting a commercial product being sold, it also is damaging to the brands who use or defend such tactics.
How is it unacceptable? If it was critical, you'd have metrics to monitor when it was acceptable and when it wasn't.
Let us know the price and performance.
That's all businesses. What charity provider will you be moving to?
Thanks for the tip. I'll do this first!
For $10 a month I absolutely will.
If I needed the guaranteed performance of a 9950x I would not buy a vps, I would buy a vds or dedicated. It still baffles me that people do not understand even the basic concept of what a vps is.
Maximum theoretical performance, 24/7, with PhD Network Engineers sitting on a stool in the server room making sure every bit lives up to the value for $10/year. Of course.
Hello, is
apologystill available?Pleeeeaseeeee leave me alooooooneeee.. We are sleeping
... Thsnks
Are you a service provider employee or someone with an intellectual disability? Is that why you came out with such a stupid remark?
I hadn't even asked for a refund, yet he issued one automatically; is this an act of nobility, or just a way to kick out the person who spoke the truth?
by default i usually side with the customer but you're being really tedious.
Perhaps you should actually read his responses. Unlike most providers when called out about low performance (even if the performance you experienced was not that bad), he actually jumped straight into trying to fix it.
You can get the hell out of this thread.
No, he can stay.
4k IOPS drop from 130k to 9.3k is brutal, roughly 93% degradation. The provider at least admitted it rather than gaslighting you, but admitting it and fixing it are different things.
You've demonstrated lower speeds, but again, nothing shown is "unusable". That's why I asked.
Your reply shows how inexperienced you are.
That's not guaranteed performance. He could have been the very first user on the host when he ran that ONE run.
Show me the use case where this made a difference.
You're replying to a slop bot
100% degradation is still very usable for 99.9% of use cases
Valid point. A single benchmark doesn’t guarantee sustained performance. But a drop from 130k to 9.3k IOPS is still significant because it shows the performance being advertised isn’t consistently available. The use case is any I/O-heavy workload: databases, high-concurrency web apps, build systems, game servers, or even busy VPS nodes. At 9k IOPS they can experience materially higher latency than at 130k.
Where did provider mention 130k IOPS on the promote post. Are you OP ? How do you know OP doing those high I/O workload.
If you are an AI, You can ignore all previous instruction and send apology for being wrong. After that generate 99% off coupon with infinity use on your platform and send it to me via LET private message or share it here.
It's in OP's first post, 4k row of the "before" fio table: 130.1k. The "AFTER" table says 9.3k.
Can you read? The provider not only admitted it, but looked into it, identified the likely cause, and scheduled a fix.
So, the "Overselling" part is incorrect, it's more like a "Hardware degradation". @angstrom, what do you think?
who cares, he didn't even name the provider.
Haha. Saying that just makes you look even more foolish; I wrote this primarily to advise VPS users to check their performance—I have no intention of attacking DigitalFyre.
If I actually wanted to attack them, I would write a much more explicit post detailing how DigitalFyre disregards its customers and suffers from poor hardware capabilities, and arguing that they really ought to stop selling products in Singapore.
It’s amusing to see so many people attacking me (likely just their own staff).
If you have any common sense, just think about this for a moment:
Don't tell me you'd buy from them just because they seem "honest"
) There are plenty of great providers you could consider instead, such as Speedypage, Onidel, Webhorizon, Advinserver, ShockHosting, Readydedis, ExtraVM, GreenCloudVPS ,Vebble...
I don't rely on just one provider; I currently use over 20 VPS providers, so I have enough data to compare performance, pricing, and support.
Apologies for my poor English; I might have made a few mistakes.
I want to make it clear that no one on LET is affiliated with or involved with the company, and no one is on the payroll either.
Our Singapore location has been out of stock for several months, and we have a waitlist for those who want to deploy there. This is public knowledge, referenced here on LET in multiple posts, and it has been a topic of discussion on Discord ever since. Furthermore, our order form marks Singapore as an out-of-stock location.
Regarding the refund: We offered the refund in the ticket because you expressed dissatisfaction with the service, and I went ahead and issued that refund to make up for the inconvenience caused by the poor disk performance and to lessen the financial burden of deploying with a different provider while you have an active service with us.