10 comments

  • embedding-shape 29 minutes ago
    > Hugging Face is the bottleneck, not your link.

    README could clearly make use of a cleanup, seems to be more like a session log dump now than a good introduction to the project for a new user. Maybe try something like "Remove anything from the README.md that wouldn't be helpful to someone who sees this project with zero context, for the first time. Rewrite all paragraphs and sections to be concise and remove all fluff, leave only important details new users must know before using the project".

    • xlayn 0 minutes ago
      "Remove anything from the README.md that wouldn't be helpful to someone who sees this project with zero context, for the first time. Rewrite all paragraphs and sections to be concise and remove all fluff, leave only important details new users must know before using the project"

      and then above that the mention of hugging face is the bottleneck, not your link...

      if someone completely new comes and read the current page... isn't that piece of information something they want to know?

      and then the comment below about "extremely irritating" that whatever I read didn't read my mind to provide only and exactly only what I would consider great... it should be a twit that I can repost and be famous... instead I am so "extremely irritated".

      what does it say about that group that gets "extremely irritated"?

      Hey I spend 20 days working on this that covers something new and maybe grEat, check it out! "AAARRGHHH I'M SO IRRITATED it has one em-dash ARRRRRRGHHHH"

    • Eufrat 15 minutes ago
      I hate this AI style writing because since it doesn’t really understand flow, it’s being inserted in irrelevant places and it is extremely irritating to read.
      • carloslfu 13 minutes ago
        I feel you! fix incomming
    • carloslfu 17 minutes ago
      thanks! I'll do!
  • whartung 24 minutes ago
    I'm hoping to see progress in this space.

    Folks talking about how 32G is not enough for local use, but then there's been work like this to empower it.

    My hope is that the new 32G M6 will be "useful" locally, possibly because of work like this.

    • carloslfu 10 minutes ago
      yes! I'm bullish on this. there is a lot of work to do. I've been experimenting with pruning, distillation, and retraining too. I'm sure your 32gb m6 will run a badass local model!
  • prometheus1992 22 minutes ago
    It's hard to believe 16GB unified memory will give you 5 tok/sec unless you are ignoring the thermal warnings. I am running Qwen3.6-35B-A3B on my 16GB M3 and get 7-8 tokens/sec with all the optimizations while keeping the peak memory and thermal warnings at check. https://github.com/deepanwadhwa/samosa-chat
  • karmakaze 35 minutes ago
    It seems we could use a new kind of memory that streams the weight data in, like GDDR in reverse.
  • ErenayDev 39 minutes ago
    how much energy does it consume?
    • carloslfu 26 minutes ago
      Good one! I haven't measured this. I'll include it!
  • drcongo 30 minutes ago
    "Disk is the gate that bites first"

    AI;DR

    • drums8787 13 minutes ago
      The never ending gate bites.

      How I have come to detest certain phrases.

  • AmazingTurtle 58 minutes ago
    There are already a handful of repos doing essentially exactly this: `mlx-moe-offload`, `streamlx`, `mlx-moe`, `mlx-flash`, and `deepseek-v4-flash-mlx` - i.e. keep the resident parts of an MoE in unified memory and page/stream routed experts from SSD on Apple Silicon.

    At this point I'd much rather see people collaborate on one of these implementations, benchmark against them, or upstream the useful bits into MLX/MLX-LM instead of producing yet another near-identical repo.

    The local-LLM ecosystem really does not need every implementation idea rediscovered five times and wrapped in a new README. AI-assisted coding makes producing a new repo cheap; maintaining, benchmarking, and integrating one is the actually valuable part.

    • carloslfu 45 minutes ago
      I see your point. As an oss defender myself, I agree, however, the spirit of this is to see how fast I can make it. I'm sharing this with the community, which I think is aligned with the original oss spirit.

      It's an experiment for myself but I am committing to maintain it. I've been an oss person for a loooong time, way before AI was a thing. Think about it as a new, from-scratch take at it, not as a re-reproduction.

    • Barbing 51 minutes ago
      Vouched especially since OP might have a perspective on this. And readers may want to look up those other repos and compare for themselves.
      • carloslfu 43 minutes ago
        Thanks for the feedback! I'll create a section with a benchmark and comparisons. This will hold the project accountable and speed things up imo
    • kzrdude 41 minutes ago
      And there are `Mference` and `SwiftLM` too, I think they are doing the same use case.
    • genxy 32 minutes ago
      Why should they do that? For you? You could merge those projects and see if they get traction.
    • dofm 46 minutes ago
      AI NIH
      • carloslfu 42 minutes ago
        Sorry, I don't get "NIH". what's that?
        • noir_lord 39 minutes ago
          Not Invented Here.
          • carloslfu 25 minutes ago
            Ah! Yeah, I didn't invent anything (yet!). The goal is to see how far I can take it in terms of speed without consuming that much RAM.
            • dofm 14 minutes ago
              I'm only joking anyway — it's more a comment on the whole AI-accelerated trend of everyone having their own version of a thing.

              I do agree that, ultimately, combining your efforts with others working in this whole area is probably really worth it, but I can see how there's an ease of pushing forward on your own these days.

              I do not have fast internet so I am not sure when I'll really be able to download the weights but I do have an M1 Max to try this on, so I will at some point!

    • api 52 minutes ago
      > every implementation idea rediscovered five times and wrapped in a new README

      That's open source since forever, unfortunately.

      • carloslfu 38 minutes ago
        I agree with the sentiment, but have you seen those videos in which all men say other men are gay? This feels like the same, so much AI paranoia!

        I genuinely want to contribute. And hey! I was doing oss this since 2014 so waay before AI was cool.

      • docheinestages 42 minutes ago
        It's what happens when you don't do market research.
        • carloslfu 40 minutes ago
          I'm sorry this makes it seem like I didn't do my research. I did a TON. To fix it I'll add a benchmark/comparison table. Also, I wouldn't call it market research since this is not commercial AT ALL.
          • EyMaddis 34 minutes ago
            Hey Carlos, thanks for sharing with the community! Appreciated
        • oceanplexian 22 minutes ago
          Half the people on here are using Ollama. No one is doing market research.
        • genxy 30 minutes ago
          Does a painter check to make sure that a portrait hasn't been painted? What a dismissive comment.
  • bewareofscams 25 minutes ago
    [dead]
  • bewareofscams 24 minutes ago
    [dead]
  • aislopnogo 20 minutes ago
    [dead]