Howdy, Stranger!

It looks like you're new here. If you want to get involved, click one of these buttons!


New on LowEndTalk? Please Register and read our Community Rules.

All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.

GPT-6 Astra

2

Comments

  • Pretty much, i dont bother with multi provider or harness either, i just plug shit into whatever IDE gives free credits at the time, else fall back to Gemini for paid/account, GLM for api/usage

    Thanked by 1rpqu
  • @SplitIce said: That sounds like as much work to implement as it would just be to do the work.

    Yes, some manual work is ofcourse but it worth for the results.

  • consumed 3b input tokens and 60m output tokens just 2 months o:)

  • AndreixAndreix Host Rep, Veteran
    edited September 5

    @William said:
    Pretty much, i dont bother with multi provider or harness either, i just plug shit into whatever IDE gives free credits at the time, else fall back to Gemini for paid/account, GLM for api/usage

    I use AI with a master brain (most of the time my self-hosted K3 or GLM5.2) and rarely Claude.
    That brain is whipping the asses of all free-to-use LLM models. Have like ~50 that I can use at no cost at all... lol.

    https://i.ibb.co/7xd4XVqQ/image.png

  • @Levi said:

    @TimboJones said:
    I'm pretty impressed by the reverse engineering of apps by AI. It's fixed a couple of stability issues the actual developer couldn't fix in over a decade. One developer I know said he was close to beating echostar ecm's with a basic GPU like 10 years ago, probably doable with a modern GPU and proper teardown of nagra protocol.

    Even you sucumb to llm? Tafak Timbo.

    I support safe and productive AI. I use it at work all the time for processing data quickly and rewriting emails and documents because me rite no gud.

    Thanked by 1forest
  • aphexaphex Member
    edited September 5

    @Andreix said: IDK why google keeps pushing -flash models instead of seriously working on a Fable/Astra/K3 rival

    One of their largest use cases is as cheap as possible bulk inference flash for the top of search result page, google assistant, hey google, nest (since users expect this usage free)

    The others companys don't have this use case or this customer base, for thing like google assistant or siri, you need LLM that answer as fast as possible while cheap, but google wants everything server sided to log the datas and apple wants fast models running local on iphone

  • @forest said:

    @shengade said:
    It's online today. If you want to use it at a low discount, you can contact me

    Huh?

    I operate a token aggregation site that sells GPT tokens at a 50% discount

  • LeviLevi Veteran

    @shengade said:

    @forest said:

    @shengade said:
    It's online today. If you want to use it at a low discount, you can contact me

    Huh?

    I operate a token aggregation site that sells GPT tokens at a 50% discount

    Nice. What methods do you use to obtain tokens?

  • @Levi said:

    @shengade said:

    @forest said:

    @shengade said:
    It's online today. If you want to use it at a low discount, you can contact me

    Huh?

    I operate a token aggregation site that sells GPT tokens at a 50% discount

    Nice. What methods do you use to obtain tokens?

    We use official discounted prices in each region and every time to purchase subscription packages。

  • LeviLevi Veteran

    @shengade said:

    @Levi said:

    @shengade said:

    @forest said:

    @shengade said:
    It's online today. If you want to use it at a low discount, you can contact me

    Huh?

    I operate a token aggregation site that sells GPT tokens at a 50% discount

    Nice. What methods do you use to obtain tokens?

    We use official discounted prices in each region and every time to purchase subscription packages。

    Do you use carded cc and stolen accounts?

  • I think my claude Opus 5 + my Gemini Pro (Free) is enough for me.
    I wanted to try Fable but leave it for now

  • @Levi said:

    @shengade said:

    @Levi said:

    @shengade said:

    @forest said:

    @shengade said:
    It's online today. If you want to use it at a low discount, you can contact me

    Huh?

    I operate a token aggregation site that sells GPT tokens at a 50% discount

    Nice. What methods do you use to obtain tokens?

    We use official discounted prices in each region and every time to purchase subscription packages。

    Do you use carded cc and stolen accounts?

    This is illegal, we use reasonable discounts

  • @shengade said: This is illegal, we use reasonable discounts

    I think they would prefer more information given openai does not really do discounts or regional pricing outside of Go

  • @SplitIce said:

    @sreekanth850 said:

    @SplitIce said: The biggest thing that burns tokens for me currently with GPT 5.6 is using AI to prototype and benchmark improvements in performance critical code.

    I had done optimization but with slight different pathway.
    use chatgpt to connect the repo, audit the code. and then do a cpu profiling and then give the report to chat, and ask to plan for any remaining optimization. Implement the optimization with luna High or terra.
    pre and post benchmark. one optimization at a time. if regressed, go back. Repeat.

    That sounds like as much work to implement as it would just be to do the work.

    He is saving tokens

  • @aphex said:

    @shengade said: This is illegal, we use reasonable discounts

    I think they would prefer more information given openai does not really do discounts or regional pricing outside of Go

    Yes, but OpenAI will provide some promotional activities for special groups

  • @aphex said: The others companys don't have this use case or this customer base,

    Meta has a similar use case but fails to use it thinking they are a coding outlet as does Apple but no real AI offer in Siri/iOS yet, Meta same as Google also offers inference basically for free.

    Gemini 3.8 Flash is IN 0.00$ and produces 2000 tok/s of not even entirely garbage for primarily the search results, it also runs on Google TPUs and not Nvidia cards which is nice to rely not on single source (Gemini = Google TPUs, Zai = Huawei, ChatGPT etc. = Nvidia, AWS = Amazon)

  • @mhpteam said:
    consumed 3b input tokens and 60m output tokens just 2 months o:)

    I do 1-2b per day

  • @vitobotta said:

    @mhpteam said:
    consumed 3b input tokens and 60m output tokens just 2 months o:)

    I do 1-2b per day

    That's not the flex you think it is. :D

  • rpqurpqu Member

    @forest said:

    @vitobotta said:

    @mhpteam said:
    consumed 3b input tokens and 60m output tokens just 2 months o:)

    I do 1-2b per day

    That's not the flex you think it is. :D

    Token cache likely included.

  • @vitobotta said:

    @mhpteam said:
    consumed 3b input tokens and 60m output tokens just 2 months o:)

    I do 1-2b per day

    drinking a lake per hour but yeah the website was beautiful

    Thanked by 1Saragoldfarb
  • Can we at least not act like we don't know how water cooling loops work? The permeation isn't that extreme

    Thanked by 2meowwcc forest
  • Youd be right if they would be using loops and often not just evaporation cooling, my work DC uses fresh water to cool the HVAC radiators because water is flat few while power is metered

  • sreekanth850sreekanth850 Member
    edited September 6

    @Anayx said:

    @SplitIce said:

    @sreekanth850 said:

    @SplitIce said: The biggest thing that burns tokens for me currently with GPT 5.6 is using AI to prototype and benchmark improvements in performance critical code.

    I had done optimization but with slight different pathway.
    use chatgpt to connect the repo, audit the code. and then do a cpu profiling and then give the report to chat, and ask to plan for any remaining optimization. Implement the optimization with luna High or terra.
    pre and post benchmark. one optimization at a time. if regressed, go back. Repeat.

    That sounds like as much work to implement as it would just be to do the work.

    He is saving tokens

    You will understand when you optimize for microseconds. especially when you work with libraries like simd json. I had ported this library to C# for a parsing project and the optimization you are doing will be for nano seconds and microseconds. That to match the C# implementation with native C++. i was close.

  • Using AI for performance prototyping and CPU profiling iteration is definitely a smart workflow, though sometimes managing token quotas feels like playing a survival game. Thanks for sharing the insights

  • 150mil GLM 5.3/Flash, Hy3, Deepseek tokens: https://kiraai.vn/dashboard/

    No daily limit it seems, no payment info, i used some trash google accounts

    Thanked by 1rpqu
  • AndreixAndreix Host Rep, Veteran

    @William said:
    150mil GLM 5.3/Flash, Hy3, Deepseek tokens: https://kiraai.vn/dashboard/

    No daily limit it seems, no payment info, i used some trash google accounts

    Shit that would fly so fast... and I bet those are total tokens, not only input.

    https://i.ibb.co/rGNfj2qb/image.png

  • JordJord Moderator, Host Rep, Veteran, Megathread Squad

    Astra is goat tier

    Thanked by 1meowwcc
  • AndreixAndreix Host Rep, Veteran

    @Jord said:
    Astra is goat tier

    Bash me, but I like GPT-5.6-Sol more. <3
    (if we stick to GPT family; else K3 FTW!)

    Thanked by 1shaewe
  • JordJord Moderator, Host Rep, Veteran, Megathread Squad

    @Andreix said:

    @Jord said:
    Astra is goat tier

    Bash me, but I like GPT-5.6-Sol more. <3
    (if we stick to GPT family; else K3 FTW!)

    You have been bashed sir, Sol is also prem, but Astra is premer

    Thanked by 1meowwcc
  • AndreixAndreix Host Rep, Veteran

    @Jord said:

    @Andreix said:

    @Jord said:
    Astra is goat tier

    Bash me, but I like GPT-5.6-Sol more. <3
    (if we stick to GPT family; else K3 FTW!)

    You have been bashed sir, Sol is also prem, but Astra is premer

    Wait until you use the Pro variant of the model. Your usage window is dropping like 50% per minute. LOL.

Sign In or Register to comment.