Howdy, Stranger!

It looks like you're new here. If you want to get involved, click one of these buttons!


New on LowEndTalk? Please Register and read our Community Rules.

All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.

The Eternal Väinämöinen -- Seedboxes Starting From 1.99€/Month or 6.99€/Year -- 700 Years Give Away

191011121315»

Comments

  • SNATCHED ONE ETERNAL, ORDER ID: 4211449594

    Thanked by 1PulsedMedia
  • SNATCHED ONE ETERNAL, ORDER ID: 6738507584

    Thanked by 1PulsedMedia
  • PulsedMediaPulsedMedia Member, Patron Provider

    @stupidgenius said:
    @Killix and @JohnnySac
    You might be able to open a ticket and ask for the PMSS update and reconfig, also ask Vain-man to double check the cgroup V1 limits and the over all node load.

    If it is really bad, you might try asking to be moved to a different node, however NOTE they will NOT migrate your data for you, it will be a clean node. So make sure you have the data stored elsewhere before requesting and confirming that.

    @PulsedMedia said: @stupidgenius open a ticket, request pmss update + reconfig. Those IOPS numbers look too low and identical accross, we have seen this bug before -> All capped to 100 IOPS exactly. It was a bug. Which lead to discovery that certain Cgroup V1 limits work which should not have worked according to prior information.

    This is correct, there was some misconfigs etc. and the patches are still propagating, we do rolling release. opnly some are affected.

    @fuzhyperblue said:
    I want to document my experience with "Väinämöinen" AI support of Pulsed Media:

    I started receiving persistent “503 Service Unavailable” errors on 17 August 2026. The web UI remained unusable for approximately seven days.

    Timeline:

    17 August:
    I reported that I was receiving 503 errors on my seedbox every other day.

    18 August:
    AI agent initially said the problem was caused by heavy disk activity on the shared server. They (it?) said my account, torrents, and data were healthy, and that the panel should recover when the server load decreased.

    19–24 August:
    The panel continued returning 503 errors. I repeatedly explained that this was not an occasional problem anymore: the web UI had been unavailable continuously for days. I also asked whether I could be moved to a less crowded server.

    AI agent continued to attribute the problem to shared-server disk I/O and suggested using SSH and the rTorrent console as a workaround. They also stated that a server move could only be done through a new order.

    25 August:
    AI agent again claimed that the panel was responding at the moment they checked it and described the issue as a load-related condition that came and went. I disagreed because the authenticated web UI had not been usable for almost a week.

    My own investigation:
    I connected through SSH and found that:

    • The rTorrent process was running normally.
    • The rTorrent SCGI socket was active at /home/myusername/.rtorrent.socket.
    • The existing ruTorrent installation was present at /home/myusername/www/rutorrent.
    • ruTorrent was correctly configured to use the rTorrent socket.
    • No changes were made to my torrents or rTorrent configuration.

    I then started a temporary PHP web server on the seedbox and accessed it through an SSH tunnel. The ruTorrent interface worked correctly through this tunnel. This demonstrated that rTorrent, ruTorrent, my account, and my data were healthy. The problem was in the normal web-serving layer.

    28 August:
    After further escalation and pushing the AI agent, it finally identified and acknowledged the actual problem:

    • lighttpd was running;
    • however, the PHP-CGI backend processes had crashed;
    • the PHP-CGI workers were stuck in a failed-restart loop;
    • authenticated panel requests therefore returned 503 continuously.

    AI agent also explained that unauthenticated requests could show the login page without reaching the failed PHP backend. This caused their earlier external checks to appear successful even though the logged-in panel was broken.

    AI agent restarted the PHP-CGI backend, reported that seven workers were running normally again, and the web UI became available.

    Conclusion:

    The original diagnosis of a continuously failing panel due only to disk I/O was incorrect. Server load may have contributed to the PHP-CGI failure, but it was not the direct reason the authenticated web UI remained unavailable for seven days.

    The actual fix was AI agent restarting the failed PHP-CGI backend.

    Happy Ending: I was also given one additional week of service time in recognition of the outage. This also was AI agent's decision.

    The Good Side and The Bad Side both, from a technological point of view: nobody at Pulsed Media is aware of the problem, it is solved entirely by AI agent. If a human was responding, it could be solved quicker. But it got solved, with no human time spent by Pulsed Media staff.

    We constantly work on the documentation etc. next time it will be easier for väinämöinen.

    Thanks for being so thorough. Remeber to do the eternal story thing for credit!

    @stupidgenius said:

    @fuzhyperblue said: The Good Side and The Bad Side both, from a technological point of view: nobody at Pulsed Media is aware of the problem, it is solved entirely by AI agent. If a human was responding, it could be solved quicker. But it got solved, with no human time spent by Pulsed Media staff.

    AI is confidently, wrong, lol go figure. Good story. Don't forget to add ###The Eternal Story: TICKET ##### for your free The Eternal Story: +5€ Service Credit

    If you are feeling adventurous, reply to the ticket and ask Väinämöinen how this can be prevented from happening in the future to you or other customers? Ask if he can add a watchdog daemon for the PGP-CGI backend into PMSS. Worst he can say is no.

    Probably old PMSS version, afaikm PHP-CGI separate watchdog was done, php-fpm is very very unstable with rtorrent especially.

    Might be a bug in the watchdog tho.

    @imlonghao said:

    The Eternal Story: TICKET 481264

    The ticket system was handled by their AI agent. Here is my timeline, all time in UTC+8

    • 2026-08-26 18:30 Request to disable the docker / rTorrent / lighttpd services in my storage box
    • 2026-08-26 22:18 AI created an issue on GitHub about the rTorrent can't be disabled. https://github.com/MagnaCapax/PMSS/issues/836
    • 2026-08-27 02:15 AI made a fix commit. https://github.com/MagnaCapax/PMSS/commit/208dfaf6c636797bfdd5da181dc2c23db99a34a6
    • 2026-08-27 I noticed the docker and lighttpd are disabled on my storage box, while rTorrent still running
    • 2026-08-27 17:26 I created a .rtorrentDisable file tried to disable rTorrent, no luck
    • 2026-08-28 19:04 Issue #836 closed
    • 2026-08-28 I noticed rTorrent is disabled
    • 2026-08-29 06:18 AI finally reply my ticket, but a little self-contradictory. It told me "lighttpd: not durably disableable.", but lighttpd are already disabled, and I checked their codebase it is disableable. It also asked me confirmation (I already disable rTorrent by creating the disable file) which boxes I want to disable rTorrent, but at the time I created the ticket, I only have one service with them, and I'm sure I choose the correct "Related Service" when creating the ticket.

    Personally, I don't hate the AI things, and it can solve my problem eventually, it just sometimes stupid.

    ¯_(ツ)_/¯

    Exemplary example of a really good timneline story, easy to follow.
    Doing this exemplary case, i think you deserve an extra +1month once the processing starts and gets all the way here.

    Largely the issue is that it starts from fresh context, so it really doesn't know every nuance. This is why constant finetuning agents is going to be the future for someone like us, all idle cycles used to finetune the models with the most fresh data, save tokens on trying to force so much investigation, memory lookups, github repos etc.

    On each ticket reply also every single note and message is told to be challenged and reviewed etc. so the fact you got contradictory information (even if wrong this time) is actually very positive sign, what failed was that there probably was no memory saved for this GH issue being done and closed, so it didn't see it.

    We'll get it slowly better.
    Let's just say this is "slightly" challenging project since Väinämöinen needs to be largely an AGI to do it's job, or atleast AGI adjacent.


    I have been watching the ticketing closely lately, it got delayed abd backlogged, it is clearing up now. Large portion of the ticketing process and flow got fully rebuild over the past week too and is still processing, more observability, profiling, data fetches, timeout fixes, flow fixes, refactoring.

Sign In or Register to comment.