Vectorized and performance-portable Quicksort (2022)

(opensource.googleblog.com)

140 points | by mococa 1 hour ago

16 comments

  • zX41ZdbW 1 hour ago
    Strange to see it here, the article is quite old.

    Since pdqsort, vqsort, and glide sort, the current state-of-the-art are driftsort and ipnsort.

    I've integrated them into ClickHouse: https://github.com/ClickHouse/ClickHouse/pull/106650

  • bee_rider 1 hour ago
    Well, it came out a while ago, so maybe we can be a bit silly:

    There’s something sort of beautiful about mergesort and heapsort. Their names tell you what their main idea is, and how they work is immediately obvious.

    Quicksort, on the other hand, has nothing beautiful about it and is named after it’s one redeeming feature (that it is quick for a lot of cases).

    • rodrigosetti 36 minutes ago
      Maybe if Hoare had called it Partitionsort in 1960, the name wouldn’t have been memorable enough to catch on and become so popular.
    • thesz 38 minutes ago
      The beauties of quicksort are that it sorts in-place and that it is embarrassingly simple.

      The in-place property can be utilized to make it very close to cache-oblivious algorithm.

    • teiferer 38 minutes ago
      How is it less beautiful?

      Honest question, curious to hear about what that means to you.

  • minitech 1 hour ago
    Actual title: “Vectorized and performance-portable Quicksort” (2022).

    Actual sense in which it’s first:

    > Happily, modern instruction sets (Arm SVE, RISC-V V, x86 AVX-512) include a special instruction suitable for partitioning. Given a separate input of yes/no values (whether an element is less than the pivot), this "compress-store" instruction stores to consecutive memory only the elements whose corresponding input is "yes". We can then logically negate the yes/no values and apply the instruction again to write the elements to the other partition. This strategy has been used in an AVX-512-specific Quicksort. But what about other instruction sets such as AVX2 that don't have compress-store? Previous work has shown how to emulate this instruction using permute instructions.

    > We build on these techniques to achieve the first vectorized Quicksort that is portable to six instruction sets across three architectures, and in fact outperforms prior architecture-specific sorts.

  • starcast2026 5 minutes ago
    I didn't like the 9 MB image file in the blog. It took me few seconds to fully render the image.
  • djsavvy 1 hour ago
    Definitely needs (2022) in the title, I was a bit confused!
  • mixologic 1 hour ago
    > Our implementation uses Highway's portable SIMD functions, so we do not have to re-implement about 3,000 lines of C++ for each platform.

    Would they do the same thing today or have an LLM re-implement those 3000 lines of c++ ?

    • atiedebee 1 hour ago
      Id say that it is a lot more likely for the in-house, at least half a decade old library to be correct and performant than 3000 lines of an LLMs mediocre regurgitation of that code.
    • glouwbug 1 hour ago
      Guys, remember when language features allowed re-usability?
      • shadowgovt 16 minutes ago
        Barely, and rarely for C++ specifically.

        I think software engineering in general is in a bit of a discoverability crisis. So many problems actually have solutions implemented... Somewhere. If you know about them. And are speaking the same vocabulary as the original implementer to realize the solution might be applicable to your problem. It's one of the reasons that jokes exist about microservice frameworks (https://www.youtube.com/watch?v=y8OnoxKotPQ) and how "We use Hadoop to store the output from our Kafka pipe, that's populated from our Traefik layer, all monitored with Grafana in front of Loki and Prometheus, of course" is a real sentence that has actual meaning and not a fever-dream.

        LLMs are actually pretty impressive at being able to pull together disparate information from various domains into one place.

  • kg 1 hour ago
    (2022)

    If you're curious why you would want a vectorized way to sort lists of numbers, one use case is building histograms - it's much easier to build a histogram if you've sorted all your samples first

    • mcdonje 1 hour ago
      >(2002)

      I wonder what apps have implemented this now that a few years have passed.

  • sciencesama 56 minutes ago
    this was made like 4 years back most of the current algorithms use this already !
  • brrrrrm 1 hour ago
    only sorts numbers? wouldn't radix be much better?
  • glouwbug 1 hour ago
    Very nice. Let's see Paul Allen's quicksort
  • rvz 57 minutes ago
    Just imagine when candidates will get asked by pre-revenue startups to implement a vectorized version of quick-sort in person in 10 mins, just for a SWE job which they do not use this themselves.

    Only the likes of MAG 7, and a couple of hedge-funds would ask to do it since this problem directly applies to them.

    But certainly not pre-revenue startups.

  • moralestapia 1 hour ago
    [flagged]
  • tucnak 1 hour ago
    [flagged]
    • Flex247A 1 hour ago
      time to log off
      • trueno 1 hour ago
        i read that and thought the same. actual based idea its even inspired me to log off
    • sigbottle 1 hour ago
      Honestly part of me feels that way about way too many things in retrospect about my own life and interests.