Skip to content

add BLAKE3 to hashlib #83479

Description

@larryhastings
BPO 39298
Nosy @malemburg, @gpshead, @larryhastings, @tiran, @mgorny, @jstasiak, @oconnor663, @corona10, @tirkarthi, @kmaork
PRs
  • bpo-39298: Add BLAKE3 bindings to hashlib. #31686
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = <Date 2022-03-23.16:53:03.784>
    created_at = <Date 2020-01-11.04:27:40.080>
    labels = ['type-feature', 'library', '3.11']
    title = 'add BLAKE3 to hashlib'
    updated_at = <Date 2022-03-24.20:30:46.143>
    user = 'https://github.com/larryhastings'

    bugs.python.org fields:

    activity = <Date 2022-03-24.20:30:46.143>
    actor = 'oconnor663'
    assignee = 'none'
    closed = True
    closed_date = <Date 2022-03-23.16:53:03.784>
    closer = 'larry'
    components = ['Library (Lib)']
    creation = <Date 2020-01-11.04:27:40.080>
    creator = 'larry'
    dependencies = []
    files = []
    hgrepos = []
    issue_num = 39298
    keywords = ['patch']
    message_count = 62.0
    messages = ['359777', '359794', '359796', '359936', '359941', '360152', '360215', '360535', '360838', '360840', '361918', '361925', '363397', '391355', '391356', '391360', '391418', '401070', '401093', '410277', '410293', '410363', '410364', '410386', '410451', '410452', '410453', '410454', '410455', '413413', '413414', '413490', '413563', '414532', '414533', '414543', '414544', '414565', '414569', '414571', '415761', '415806', '415823', '415842', '415846', '415847', '415863', '415885', '415888', '415889', '415894', '415895', '415897', '415898', '415900', '415902', '415920', '415948', '415951', '415952', '415960', '415973']
    nosy_count = 11.0
    nosy_names = ['lemburg', 'gregory.p.smith', 'larry', 'christian.heimes', 'mgorny', "Zooko.Wilcox-O'Hearn", 'jstasiak', 'oconnor663', 'corona10', 'xtreak', 'kmaork']
    pr_nums = ['31686']
    priority = 'normal'
    resolution = 'rejected'
    stage = 'resolved'
    status = 'closed'
    superseder = None
    type = 'enhancement'
    url = 'https://bugs.python.org/issue39298'
    versions = ['Python 3.11']

    Activity

    1. larryhastings commented on Jan 11, 2020

      @larryhastings
      ContributorAuthor

      From 3/4 of the team that brought you BLAKE2, now comes... BLAKE3!

      https://github.com/BLAKE3-team/BLAKE3

      BLAKE3 is a brand new hashing function. It's fast, it's paralellizeable, and unlike BLAKE2 there's only one variant.

      I've experimented with it a little. On my laptop (2018 Intel i7 64-bit), the portable implementation is kind of middle-of-the-pack, but with AVX2 enabled it's second only to the "Haswell" build of KangarooTwelve. On a 32-bit ARMv7 machine the results are more impressive--the portable implementation is neck-and-neck with MD4, and with NEON enabled it's definitely the fastest hash function I tested. These tests are all single-threaded and eliminate I/O overhead.

      The above Github repo has a reference implementation in C which includes Intel and ARM SIMD drivers. Unsurprisingly, the interface looks roughly the same as the BLAKE2 interface(s), so if you took the existing BLAKE2 module and s/blake2b/blake3/ you'd be nearly done. Not quite as close as blake2b and blake2s though ;-)

    2. added
      stdlibStandard Library Python modules in the Lib/ directory
      type-featureA feature request or enhancement
      on Jan 11, 2020
    3. tiran commented on Jan 11, 2020

      @tiran
      Member

      I've been playing with the new algorithm, too. Pretty impressive!

      Let's give the reference implementation a while to stabilize. The code has comments like: "This is only for benchmarking. The guy who wrote this file hasn't touched C since college. Please don't use this code in production."

    4. larryhastings commented on Jan 11, 2020

      @larryhastings
      ContributorAuthor

      For what it's worth, I spent some time producing clean benchmarks. All these were run on the same laptop, and all pre-load the same file (406668786 bytes) and run one update() on the whole thing to minimize overhead. K12 and BLAKE3 are using a hand-written C driver, and compiled with both gcc and clang; all the rest of the algorithms are from hashlib.new, python3 configured with --enable-optimizations and compiled with gcc. K12 and BLAKE3 support several SIMD extensions; this laptop only has AVX2 (no AVX512). All these numbers are the best of 3. All tests were run in a single thread.

      -----------------+----------+----------+----+-----------------------
      hash algorithm|elapsed s |mb/sec |size|hash
      -----------------+----------+----------+----+-----------------------
      K12-Haswell 0.176949 2298224495 64 24693954fa0dfb059f99...
      K12-Haswell-clang 0.181968 2234841926 64 24693954fa0dfb059f99...
      BLAKE3-AVX2-clang 0.250482 1623547723 64 30149a073eab69f76583...
      BLAKE3-AVX2 0.256845 1583326242 64 30149a073eab69f76583...
      md4 0.37684668 1079135924 32 d8a66422a4f0ae430317...
      sha1 0.46739069 870083193 40 a7488d7045591450ded9...
      K12-clang 0.498058 816509323 64 24693954fa0dfb059f99...
      BLAKE3 0.561470 724292378 64 30149a073eab69f76583...
      K12 0.569490 714093306 64 24693954fa0dfb059f99...
      BLAKE3-clang 0.573743 708800001 64 30149a073eab69f76583...
      blake2b 0.58276098 697831191 128 809ca44337af39792f8f...
      md5 0.59936016 678504863 32 306d7de4d1622384b976...
      sha384 0.64208886 633352818 96 b107ce5d086e9757efa7...
      sha512_224 0.66094102 615287556 56 90931762b9e553bd07f3...
      sha512_256 0.66465768 611846969 64 27b03aacdfbde1c2628e...
      sha512 0.6776549 600111921 128 f0af29e2019a6094365b...
      blake2s 0.86828375 468359318 64 02bee0661cd88aa2be15...
      sha256 0.97720436 416155312 64 48b5243cfcd90d84cd3f...
      sha224 1.0255457 396538907 56 10fb56b87724d59761c6...
      shake_128 1.0895037 373260576 32 2ec12727ac9d59c2e842...
      md5-sha1 1.1171806 364013470 72 306d7de4d1622384b976...
      sha3_224 1.2059123 337229156 56 93eaf083ca3a9b348e14...
      shake_256 1.3039152 311882857 64 b92538fd701791db8c1b...
      sha3_256 1.3417314 303092540 64 69354bf585f21c567f1e...
      ripemd160 1.4846368 273918025 40 30f2fe48fec404990264...
      sha3_384 1.7710776 229616579 96 61af0469534633003d3b...
      sm3 1.8384831 221198006 64 1075d29c75b06cb0af3e...
      sha3_512 2.4839673 163717444 128 c7c250e79844d8dc856e...

      If I can't have BLAKE3, I'm definitely switching to BLAKE2 ;-)

    5. oconnor663 commented on Jan 13, 2020

      oconnor663mannequin
      Mannequin

      I'm in the middle of adding some Rust bindings to the C implementation in github.com/BLAKE3-team/BLAKE3, so that cargo test and cargo bench can cover both. Once that's done, I'll follow up with benchmark numbers from my laptop (Kaby Lake i5-8250U, also AVX2 with no AVX-512). For benchmark numbers with AVX-512 support, see the Performance section of the BLAKE3 paper (https://github.com/BLAKE3-team/BLAKE3-specs/blob/master/blake3.pdf). Larry, what processor did you run your benchmarks on?

      Also, is there anything currently in CPython that does dispatch based on runtime CPU feature detection? Is this something that BLAKE3 should do for itself, or is there existing machinery that we'd want to integrate with?

    6. larryhastings commented on Jan 13, 2020

      @larryhastings
      ContributorAuthor

      According to my order details it is a "8th Generation Intel Core i7-8650U".

    7. oconnor663 commented on Jan 16, 2020

      oconnor663mannequin
      Mannequin

      Ok, I've added Rust bindings to the BLAKE3 C implementation, so that I can benchmark it in a vaguely consistent way. My laptop is an i5-8250U, which should be very similar to yours. (Both are "Kaby Lake Refresh".) My end result do look similar to yours with TurboBoost on, but pretty different with TurboBoost off:

      with TurboBoost on
      ------------------
      K12 GCC | 2159 MB/s
      BLAKE3 Rust | 1787 MB/s
      BLAKE3 C Clang | 1588 MB/s
      BLAKE3 C GCC | 1453 MB/s

      with TurboBoost off
      -------------------
      BLAKE3 Rust | 1288 MB/s
      K12 GCC | 1060 MB/s
      BLAKE3 C Clang | 1094 MB/s
      BLAKE3 C GCC | 943 MB/s

      The difference seems to be that with TurboBoost on, the BLAKE3 benchmarks have my CPU sitting around 2.4 GHz, while for the K12 benchmarks it's more like 2.9 GHz. With TurboBoost off, both benchmarks run at 1.6 GHz, and BLAKE3 does better. I'm not sure what causes that frequency difference. Perhaps some high-power instruction that the BLAKE3 implementation is emitting?

      To reproduce these numbers you can clone these two repos (the latter is where I happen to have a K12 benchmark):

      https://github.com/BLAKE3-team/BLAKE3
      https://github.com/oconnor663/blake2_simd

      Then in both cases checkout the "bench_406668786" branch, where I've put some benchmarks with the same input length you used.

      For Rust BLAKE3, at the root of the BLAKE3 repo, run: cargo +nightly bench 406668786

      For C BLAKE3, the command is the same, but run it in the "./c/blake3_c_rust_bindings" directory. The build defaults to GCC, and you can "export CC=clang" to switch it.

      For my K12 benchmark, at the root of the blake2_simd repo, run: cargo +nightly bench --features=kangarootwelve 406668786

    8. oconnor663 commented on Jan 17, 2020

      oconnor663mannequin
      Mannequin

      I plan to bring the C code up to speed with the Rust code this week. As part of that, I'll probably remove comments like the one above :) Otherwise, is there anything else we can do on our end to help with this?

    9. 53 remaining items

    10. malemburg commented on Mar 23, 2022

      @malemburg
      Member

      On 23.03.2022 17:53, Larry Hastings wrote:

      Ok, I give up.

      Sorry to spoil the fun, but there's no need to throw
      everything in the bin ;-)

      A lean and fast blake3 C package would still be a great thing
      to have on PyPI, e.g. provide support for platforms, which
      Jack's blake3 Rust package doesn't cover, e.g.

      Raspis:
      https://www.piwheels.org/project/blake3/

      Android (e.g. via termux):
      https://wiki.termux.com/wiki/Main_Page
      https://wiki.termux.com/wiki/Python

      etc.

    11. larryhastings commented on Mar 23, 2022

      @larryhastings
      ContributorAuthor

      The Rust version is already quite "lean". And it can be much faster than the C version, because it supports internal multithreading. Even without multithreading I bet it's at least a hair faster.

      Also, Jack has independently written a Python package based around the C version:

      https://github.com/oconnor663/blake3-py/tree/master/c_impl

      so my making one would be redundant.

      I have no interest in building standalone BLAKE3 PyPI packages for Raspberry Pi or Android. My goal was for BLAKE3 to be one of the "included batteries" in Python--which would have meant it would, eventually, be available on the Raspberry Pi and Android builds that way.

    12. malemburg commented on Mar 23, 2022

      @malemburg
      Member

      With "lean" I meant: doesn't use much code and is easy to compile
      and install.

      I built a wheel from Jack's experimental package and it comes out to
      just under 100kB on Linux x64, compared to around the 1.1MB the
      Rust wheel needs:

      Archive: blake3_experimental_c-0.0.1-cp310-cp310-linux_x86_64.whl
      Length Date Time Name
      --------- ---------- ----- ----
      348528 2022-03-23 18:38 blake3.cpython-310-x86_64-linux-gnu.so
      3183 2022-03-23 18:38 blake3_experimental_c-0.0.1.dist-info/METADATA
      105 2022-03-23 18:38 blake3_experimental_c-0.0.1.dist-info/WHEEL
      7 2022-03-23 18:38 blake3_experimental_c-0.0.1.dist-info/top_level.txt
      451 2022-03-23 18:38 blake3_experimental_c-0.0.1.dist-info/RECORD
      --------- -------
      352274 5 files

      Archive: blake3-0.3.1-cp310-cp310-manylinux_2_5_x86_64.manylinux1_x86_64.whl
      Length Date Time Name
      --------- ---------- ----- ----
      3800 2022-01-13 01:26 blake3-0.3.1.dist-info/METADATA
      133 2022-01-13 01:26 blake3-0.3.1.dist-info/WHEEL
      48 2022-01-13 01:26 blake3/init.py
      4195392 2022-01-13 01:26 blake3/blake3.cpython-310-x86_64-linux-gnu.so
      382 2022-01-13 01:26 blake3-0.3.1.dist-info/RECORD
      --------- -------
      4199755 5 files

      I don't know why there is such a significant difference in size. Perhaps
      the Rust version includes multiple variants for different CPU
      optimizations ?!

    13. larryhastings commented on Mar 23, 2022

      @larryhastings
      ContributorAuthor

      I can't answer why the Rust one is so much larger--that's a question for Jack. But the blake3-py you built might (should?) have support for SIMD extensions. See the setup.py for how that works; it appears to at least try to use the SIMD extensions on x86 POSIX (32- and 64-bit), x86_64 Windows, and 64-bit ARM POSIX.

      If you were really curious, you could run some quick benchmarks, then hack your local setup.py to not attempt adding support for those (see "portable code only" in setup.py) and do a build, and run your benchmarks again. If BLAKE3 got a lot slower, yup, you (initially) built it with SIMD extension support.

    14. gpshead commented on Mar 23, 2022

      @gpshead
      Member

      To anyone else who comes along with motivation:

      I'm fine with blake3 being in hashlib, but I don't want us to guarantee it by carrying the implementation of the algorithm in the CPython codebase itself unless it gains wide industry standard-like adoption status.

      We should feel free to link to both the Rust blake3 and C blake3-py packages from the hashlib docs regardless.

    15. gpshead commented on Mar 23, 2022

      @gpshead
      Member

      Performance wise... The SHA series have hardware acceleration on modern CPUs and SoCs. External libraries such as OpenSSL are in a position to provide implementations that make use of that. Same with the Linux Kernel CryptoAPI (https://bugs.python.org/issue47102).

      Hardware accelerated SHAs are likely faster than blake3 single core. And certainly more efficient in terms of watt-secs/byte.

    16. malemburg commented on Mar 23, 2022

      @malemburg
      Member

      Here's a wheel which only includes the portable code (I disabled
      all the special cases as you suggested).

      Archive: dist/blake3_experimental_c-0.0.1-cp310-cp310-linux_x86_64.whl
      Length Date Time Name
      --------- ---------- ----- ----
      297680 2022-03-23 19:26 blake3.cpython-310-x86_64-linux-gnu.so
      3183 2022-03-23 19:26 blake3_experimental_c-0.0.1.dist-info/METADATA
      105 2022-03-23 19:26 blake3_experimental_c-0.0.1.dist-info/WHEEL
      7 2022-03-23 19:26 blake3_experimental_c-0.0.1.dist-info/top_level.txt
      451 2022-03-23 19:26 blake3_experimental_c-0.0.1.dist-info/RECORD
      --------- -------
      301426 5 files

      I didn't run any benchmarks, but it's clear that the SIMD code was
      used in my initial build and this adds some 50kB to the .so file.
      This is on a older Linux x64 box with Intel i7-4770k CPU.

      Could be that the Rust version adds several such SIMD variants and
      then branches based on the platform running the code.

      In any case, the C extension is indeed very easy to build and
      install with a standard compiler setup.

    17. gpshead commented on Mar 23, 2022

      @gpshead
      Member

      Rust based anything comes with a baseline level of Rust code overhead. https://stackoverflow.com/questions/29008127/why-are-rust-executables-so-huge

      That seems expected.

    18. larryhastings commented on Mar 24, 2022

      @larryhastings
      ContributorAuthor

      Performance wise... The SHA series have hardware acceleration on
      modern CPUs and SoCs. External libraries such as OpenSSL are in a
      position to provide implementations that make use of that. Same with
      the Linux Kernel CryptoAPI (https://bugs.python.org/issue47102).

      Hardware accelerated SHAs are likely faster than blake3 single core.
      And certainly more efficient in terms of watt-secs/byte.

      I don't know if OpenSSL currently uses the Intel SHA1 extensions.
      A quick google suggests they added support in 2017. And:

      • I'm using a recent CPU that AFAICT supports those extensions.
        (AMD 5950X)
      • My Python build with BLAKE3 support is using the OpenSSL implementation
        of SHA1 (_hashlib.openssl_sha1), which I believe is using the OpenSSL
        provided by the OS. (I haven't built my own OpenSSL or anything.)
      • I'm using a recent operating system release (Pop!_OS 21.10), which
        currently has OpenSSL version 1.1.1l-1ubuntu1.1 installed.
      • My Python build with BLAKE3 doesn't support multithreaded hashing.
      • In that Python build, BLAKE3 is roughly twice as fast as SHA1 for
        non-trivial workloads.
    19. oconnor663 commented on Mar 24, 2022

      oconnor663mannequin
      Mannequin

      Hardware accelerated SHAs are likely faster than blake3 single core.

      Surprisingly, they're not. Here's a quick measurement on my recent ThinkPad laptop (64 KiB of input, single-threaded, TurboBoost left on), which supports both AVX-512 and the SHA extensions:

      OpenSSL SHA-256: 1816 MB/s
      OpenSSL SHA-1: 2103 MB/s
      BLAKE3 SSE2: 2109 MB/s
      BLAKE3 SSE4.1: 2474 MB/s
      BLAKE3 AVX2: 4898 MB/s
      BLAKE3 AVX-512: 8754 MB/s

      The main reason SHA-1 and SHA-256 don't do better is that they're fundamentally serial algorithms. Hardware acceleration can speed up a single instance of their compression functions, but there's just no way for it to run more than one instance per message at a time. In contrast, AES-CTR can easily parallelize its blocks, and hardware accelerated AES does beat BLAKE3.

      And certainly more efficient in terms of watt-secs/byte.

      I don't have any experience measuring power myself, so take this with a grain of salt: I think the difference in throughput shown above is large enough that, even accounting for the famously high power draw of AVX-512, BLAKE3 comes out ahead in terms of energy/byte. Probably not on ARM though.

    20. tiran commented on Mar 24, 2022

      @tiran
      Member

      sha1 should be considered broken anyway and sha256 does not perform well on 64bit systems. Truncated sha512 (sha512-256) typically performs 40% faster than sha256 on X86_64. It should get you close to the performance of BLAKE3 SSE4.1 on your system.

    21. oconnor663 commented on Mar 24, 2022

      oconnor663mannequin
      Mannequin

      Truncated sha512 (sha512-256) typically performs 40% faster than sha256 on X86_64.

      Without hardware acceleration, yes. But because SHA-NI includes only SHA-1 and SHA-256, and not SHA-512, it's no longer a level playing field. OpenSSL's SHA-512 and SHA-512/256 both get about 797 MB/s on my machine.

    22. gpshead commented on Mar 24, 2022

      @gpshead
      Member

      You missed the key "And certainly more efficient in terms of watt-secs/byte" part.

    23. oconnor663 commented on Mar 24, 2022

      oconnor663mannequin
      Mannequin

      I did reply to that point above with some baseless speculation, but now I can back up my baseless speculation with unscientific data :)

      https://gist.github.com/oconnor663/aed7016c9dbe5507510fc50faceaaa07

      According to whatever powerstat -R measures on my laptop, running hardware-accelerated SHA-256 in a loop for a minute or so takes 26.86 Watts on average. Doing the same with AVX-512 BLAKE3 takes 29.53 Watts, 10% more. Factoring in the 4.69x difference in throughput reported by those loops, the overall energy/byte for BLAKE3 is 4.27x lower than SHA-256. This is my first time running a power benchmark, so if this sounds implausible hopefully someone can catch my mistakes.

    24. transferred this issue fromon Apr 10, 2022
    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    No one assigned

      Labels

      3.11only security fixesstdlibStandard Library Python modules in the Lib/ directorytype-featureA feature request or enhancement

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions