Howdy, Stranger!

It looks like you're new here. If you want to get involved, click one of these buttons!


New on LowEndTalk? Please Register and read our Community Rules.

All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.

Community GPU servers

2»

Comments

  • NeoonNeoon Community Contributor, Veteran

    @averagedatahoarder said:
    Need to bring myself to drive an hour to the datacenter this weekend lol. If it had to be one, would folks rather have Qwen3.8-27B Abliterated or something like Minimax H3?

    From past experience running videogen endpoints I feel like hosting H3 would get messy fast (gooner bait + potentially illegal content), so quite a headache to handle it all

    To be honest, I would like to try a few models, I don't even know how much Memory the system has, but Qwen 3.8 27B Abiliterated sounds good.
    You can manage then endpoint if you want, but you could also give a trusted user SSH, could compile llama.cpp from source and see what fits best.

    It would run behind a proxy anyway, otherwise we couldn't give out keys.

    Thanked by 1barbarza
  • @Neoon said:

    @averagedatahoarder said:
    Need to bring myself to drive an hour to the datacenter this weekend lol. If it had to be one, would folks rather have Qwen3.8-27B Abliterated or something like Minimax H3?

    From past experience running videogen endpoints I feel like hosting H3 would get messy fast (gooner bait + potentially illegal content), so quite a headache to handle it all

    To be honest, I would like to try a few models, I don't even know how much Memory the system has, but Qwen 3.8 27B Abiliterated sounds good.
    You can manage then endpoint if you want, but you could also give a trusted user SSH, could compile llama.cpp from source and see what fits best.

    It would run behind a proxy anyway, otherwise we couldn't give out keys.

    Honestly, I don't think llama.cpp would be best here, llama.cpp is great for lowend GPUs and offloading but vLLM/SGLang would be fastest here. It has 64GB of VRAM, could run FP8 or NVFP4. Maybe even fit some quantized version of H3 too, could have on GPU for LLMs and one for videogen. But again, annoying to handle content on videogen.

    I could consider giving you SSH but you'd have to promise not to do anything funny on it... have been burned too many times giving folks access to servers (both in cloud and "home"(colo)lab, lol)

    Thanked by 1barbarza
  • And yeah obv. a lightweight proxy on top with limiting + keys. Wish LET has OAuth, smh

    Thanked by 1barbarza
  • I wonder if I would need a patron provider tag if I gave away API keys to LET members for free for a hobby endpoint..

  • NeoonNeoon Community Contributor, Veteran

    @averagedatahoarder said:
    And yeah obv. a lightweight proxy on top with limiting + keys. Wish LET has OAuth, smh

    I already coded automatic verification for LET, people could request keys automatically with certain requirements. If you really plan to do this, this weekend, I could prepare stuff.

    @averagedatahoarder said:
    I wonder if I would need a patron provider tag if I gave away API keys to LET members for free for a hobby endpoint..

    If its free, it should be fine, but asking is always good.

  • Well, I can try but I have been really busy lately, so I can't make any guarantees. I will follow up by DM when server is back up. If you already created a LET bot for verification that's great, when I bring server back up I can take care of setting up inference, etc

  • @Obelous said:
    Are you gonna contribute? Or do you just want to leech?

    lol

  • NeoonNeoon Community Contributor, Veteran

    @averagedatahoarder said:

    @Neoon said:

    @averagedatahoarder said:
    Need to bring myself to drive an hour to the datacenter this weekend lol. If it had to be one, would folks rather have Qwen3.8-27B Abliterated or something like Minimax H3?

    From past experience running videogen endpoints I feel like hosting H3 would get messy fast (gooner bait + potentially illegal content), so quite a headache to handle it all

    To be honest, I would like to try a few models, I don't even know how much Memory the system has, but Qwen 3.8 27B Abiliterated sounds good.
    You can manage then endpoint if you want, but you could also give a trusted user SSH, could compile llama.cpp from source and see what fits best.

    It would run behind a proxy anyway, otherwise we couldn't give out keys.

    Honestly, I don't think llama.cpp would be best here, llama.cpp is great for lowend GPUs and offloading but vLLM/SGLang would be fastest here. It has 64GB of VRAM, could run FP8 or NVFP4. Maybe even fit some quantized version of H3 too, could have on GPU for LLMs and one for videogen. But again, annoying to handle content on videogen.

    I could consider giving you SSH but you'd have to promise not to do anything funny on it... have been burned too many times giving folks access to servers (both in cloud and "home"(colo)lab, lol)

    I would have to look at the benchmarks, there are some custom build engines for the 3090 and 4090, no idea about the 5090 though.
    RAM would be also interesting, turns out MoE learns to fly with fast VRAM and lot of RAM.

    I wouldn't worry about that, I run a few projects, which are sponsored (in my signature)

  • True! But I think vLLM/SGLang is about as optimized as it gets

  • @averagedatahoarder said:
    Need to bring myself to drive an hour to the datacenter this weekend lol. If it had to be one, would folks rather have Qwen3.8-27B Abliterated or something like Minimax H3?

    From past experience running videogen endpoints I feel like hosting H3 would get messy fast (gooner bait + potentially illegal content), so quite a headache to handle it all

    Both 😜

    Ideally this project would help make models available which aren’t available from other providers, such as Obliterated and H3 etc…

  • @barbarza said:

    @averagedatahoarder said:
    Need to bring myself to drive an hour to the datacenter this weekend lol. If it had to be one, would folks rather have Qwen3.8-27B Abliterated or something like Minimax H3?

    From past experience running videogen endpoints I feel like hosting H3 would get messy fast (gooner bait + potentially illegal content), so quite a headache to handle it all

    Both 😜

    Ideally this project would help make models available which aren’t available from other providers, such as Obliterated and H3 etc…

    Abliterated Qwen3 why not :)

    Abliterated H3 you're giving me PTSD from when I used to run a public videogen endpoint fuck no 😭

    Thanked by 1barbarza
  • Maybe H3 with safeguards... :p

    Thanked by 1barbarza
  • Thank you @averagedatahoarder for helping to edge this dream towards reality

    Thanked by 1Xrmaddness
  • @averagedatahoarder said:

    @barbarza said:

    @averagedatahoarder said:
    Need to bring myself to drive an hour to the datacenter this weekend lol. If it had to be one, would folks rather have Qwen3.8-27B Abliterated or something like Minimax H3?

    From past experience running videogen endpoints I feel like hosting H3 would get messy fast (gooner bait + potentially illegal content), so quite a headache to handle it all

    Both 😜

    Ideally this project would help make models available which aren’t available from other providers, such as Obliterated and H3 etc…

    Abliterated Qwen3 why not :)

    Abliterated H3 you're giving me PTSD from when I used to run a public videogen endpoint fuck no 😭

    💀

    I don’t dare ask…

  • @barbarza said:

    @averagedatahoarder said:

    @barbarza said:

    @averagedatahoarder said:
    Need to bring myself to drive an hour to the datacenter this weekend lol. If it had to be one, would folks rather have Qwen3.8-27B Abliterated or something like Minimax H3?

    From past experience running videogen endpoints I feel like hosting H3 would get messy fast (gooner bait + potentially illegal content), so quite a headache to handle it all

    Both 😜

    Ideally this project would help make models available which aren’t available from other providers, such as Obliterated and H3 etc…

    Abliterated Qwen3 why not :)

    Abliterated H3 you're giving me PTSD from when I used to run a public videogen endpoint fuck no 😭

    💀

    I don’t dare ask…

    Well, maybe LET community will be different 🙃

    Thanked by 1barbarza
  • @averagedatahoarder said: would folks rather have Qwen3.8-27B Abliterated

    I run this:
    https://huggingface.co/OBLITERATUS/Qwen3.8-27B-OBLITERATED

    Comfortable 30tok/s on M2, i can fit larger but the concurrent performance is not there.

    Thanked by 2barbarza pyrolad
  • @William said:

    @averagedatahoarder said: would folks rather have Qwen3.8-27B Abliterated

    I run this:
    https://huggingface.co/OBLITERATUS/Qwen3.8-27B-OBLITERATED

    Comfortable 30tok/s on M2, i can fit larger but the concurrent performance is not there.

    Ah this is the one by the Pliny guy, he is a grifter. Would not trust his models 😂

  • He is good at hyping his stuff up though

  • @averagedatahoarder said: Ah this is the one by the Pliny guy, he is a grifter. Would not trust his models 😂

    You just described every ai influencer in existence

  • @aphex said:

    @averagedatahoarder said: Ah this is the one by the Pliny guy, he is a grifter. Would not trust his models 😂

    You just described every ai influencer in existence

    😂

  • WilliamWilliam Veteran
    edited September 11

    @averagedatahoarder said: Ah this is the one by the Pliny guy, he is a grifter. Would not trust his models 😂

    I search models by "uncensored" and similar tags, i dont really look who provides them - for my use the more shady the better anyway

    Personally i wait now for a Deepseek 4.1 flash model, i already run the official at work but i cant exactly plug my porn bots into that cluster (well, i can, but my boss will be not amused and bill me the tokens)

  • D> @William said:

    @averagedatahoarder said: Ah this is the one by the Pliny guy, he is a grifter. Would not trust his models 😂

    I search models by "uncensored" and similar tags, i dont really look who provides them - for my use the more shady the better anyway

    Personally i wait now for a Deepseek 4.1 flash model, i already run the official at work but i cant exactly plug my porn bots into that cluster (well, i can, but my boss will be not amused and bill me the tokens)

    No way are we running DSv4.1 flash, it's twice as large (total not active) as DSv4 flash and even DSv4 flash would be out of reach

  • WilliamWilliam Veteran
    edited September 12

    yea only with SSD offloading on enough unified memory or with a 100GB card

  • Tldr went irl to day but server issues, brought server home and need to work on it a bit. Hopefully by EOW if i have time

Sign In or Register to comment.