diff --git a/CHANGELOG.md b/CHANGELOG.md
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -1,5 +1,52 @@
 # CHANGELOG for crypton
 
+## 2.1.6
+
+Two things a caller could walk into, the licence field saying what the tree
+actually holds, and the C building without a warning.
+
+* fix(pbkdf2): an output length of zero no longer takes the process down.
+  `tryFastPBKDF2_*` passed it through to C, where `assert(out && nout)`
+  aborted -- crypton's C is built without `NDEBUG`, so its assertions are
+  live in a release.  `Crypto.KDF.PBKDF2.tryGenerate` has always answered
+  with an empty result for the same request, and the fast paths now agree
+* fix(p256): `Crypto.PubKey.ECC.P256.scalarInv` returns on a zero scalar
+  rather than looping for ever.  `scalarFromBinary` accepts any 256 bits, so
+  a zero scalar is easy to come by, and the binary extended Euclid behind
+  `scalarInv` has no exit for it: zero stays even and is halved for ever.
+  A hang inside a foreign call is not interruptible, so `System.Timeout` was
+  no help either.  It now answers zero, which is what the function already
+  answered for the other input with no inverse, and what `scalarInvSafe`
+  answers for both.  crypton's own ECDSA was never exposed: it rejects a
+  zero scalar before inverting, and uses `scalarInvSafe`
+* doc(cabal): the `license:` field says what the tree holds --
+  `BSD-3-Clause AND MIT AND ISC` -- and `license-files:` lists the five
+  licence texts, where a tool looking for licences will find them.  The
+  parts of `cbits/aes/gcm_fused_x86.c` that follow picotls's `fusion` now
+  carry its MIT notice beside the file.  **Nothing is required of a user
+  that was not required before**: crypton's own code is BSD-3-Clause as it
+  always was, and the MIT and ISC code was already in the tree -- the field
+  was silent about it.  Raised by Joey Hess in #232, and settled with the
+  help of Kazuho Oku, who divided `fusion` between what derives from
+  OpenSSL and what does not, and rewrote the former upstream
+* fix(c): the sanitizer build is quiet again.  `crypton_sha256_finalize` and
+  `crypton_sha512_finalize` say that their pointers are never null, which
+  they never were.  gcc's `-Wstringop-overflow` had been reporting them as
+  writing "into a region of size 0" at "address zero" under
+  `-fsanitize=undefined`: UndefinedBehaviorSanitizer inserts a null check
+  before `memcpy`, because glibc declares `memcpy` nonnull, and the check
+  puts a null path in front of the warning pass.  Stating the contract
+  removes the path rather than the warning, and costs nothing -- compiled as
+  the package compiles it, the assembly is identical either way
+* fix(pbkdf2): an instantiation whose digest is larger than its block is
+  refused where it is written rather than after it has overflowed.  The
+  macro that builds the three PBKDF2 variants shortens a long key by hashing
+  it into a buffer the size of the block, and asserted afterwards that the
+  result fitted -- afterwards being too late, since the write has already
+  happened.  The check is a compile-time one now.  The three that exist are
+  unaffected: SHA-1, SHA-256 and SHA-512 have digests of 20, 32 and 64 bytes
+  against blocks of 64, 64 and 128
+
 ## 2.1.5
 
 crypton 2.1.3 and 2.1.4 cannot be built with GCC 14 or newer; it was
diff --git a/LICENSE b/LICENSE
--- a/LICENSE
+++ b/LICENSE
@@ -1,4 +1,5 @@
 Copyright (c) 2006-2015 Vincent Hanquez <vincent@snarc.org>
+Copyright (c) 2023-2026 Kazu Yamamoto <kazu@iij.ad.jp>
 
 All rights reserved.
 
diff --git a/README.md b/README.md
--- a/README.md
+++ b/README.md
@@ -95,31 +95,31 @@
 
 Throughput in MB/s, **higher is better**:
 
-| | crypton 1.1.5 | crypton 2.1.4 | OpenSSL | 2.1.4 / OpenSSL |
+| | crypton 1.1.5 | crypton 2.1.5 | OpenSSL | 2.1.5 / OpenSSL |
 | --- | ---: | ---: | ---: | ---: |
-| AES-128-GCM | 1335 | **6038** | 4070 | 1.48 |
-| AES-256-GCM | 1093 | **5461** | 3774 | 1.45 |
-| ChaCha20-Poly1305 | 399 | 2195 | 2213 | 0.99 |
-| SHA-1 | 729 | 1679 | 1682 | 1.00 |
-| SHA-256 | 290 | 1585 | 1571 | 1.01 |
-| SHA-512 | 449 | 770 | 750 | 1.03 |
-| SHA3-256 | 109 | 422 | 425 | 0.99 |
+| AES-128-GCM | 1362 | **6038** | 4055 | 1.49 |
+| AES-256-GCM | 1093 | **5462** | 3770 | 1.45 |
+| ChaCha20-Poly1305 | 399 | 2211 | 2229 | 0.99 |
+| SHA-1 | 727 | 1678 | 1673 | 1.00 |
+| SHA-256 | 290 | 1585 | 1579 | 1.00 |
+| SHA-512 | 463 | 804 | 751 | 1.07 |
+| SHA3-256 | 109 | 424 | 425 | 1.00 |
 
 Time per operation in microseconds, **lower is better**:
 
-| | crypton 1.1.5 | crypton 2.1.4 | OpenSSL | OpenSSL / 2.1.4 |
+| | crypton 1.1.5 | crypton 2.1.5 | OpenSSL | OpenSSL / 2.1.5 |
 | --- | ---: | ---: | ---: | ---: |
-| X25519 | 43.89 | 28.47 | 36.61 | 1.29 |
-| ECDH P-256 | 164.7 | 50.34 | 51.66 | 1.03 |
-| ECDH P-384 | 2237 | **163.4** | 835.4 | 5.11 |
-| Ed25519 sign | 28.90 | 18.69 | 33.68 | 1.80 |
-| Ed25519 verify | 47.81 | 47.77 | 111.2 | 2.33 |
-| ECDSA P-256 sign | 81.05 | 18.73 | 21.88 | 1.17 |
-| ECDSA P-256 verify | 232.6 | 70.27 | 67.51 | 0.96 |
-| ECDSA P-384 sign | 2260 | **302.9** | 879.1 | 2.90 |
-| ECDSA P-384 verify | 2668 | **467.8** | 725.2 | 1.55 |
-| RSA-2048 sign/decrypt | 758.5 | 611.2 | 659.0 | 1.08 |
-| RSA-2048 verify/encrypt | 32.79 | 30.01 | 18.85 | 0.63 |
+| X25519 | 45.31 | 28.41 | 36.48 | 1.28 |
+| ECDH P-256 | 165.4 | 51.14 | 51.65 | 1.01 |
+| ECDH P-384 | 2278 | **165.0** | 847.5 | 5.13 |
+| Ed25519 sign | 30.03 | 18.62 | 33.71 | 1.81 |
+| Ed25519 verify | 48.05 | 47.66 | 110.6 | 2.32 |
+| ECDSA P-256 sign | 81.70 | 18.96 | 21.87 | 1.15 |
+| ECDSA P-256 verify | 233.3 | 70.47 | 67.52 | 0.96 |
+| ECDSA P-384 sign | 2264 | **303.4** | 890.1 | 2.93 |
+| ECDSA P-384 verify | 2676 | **471.4** | 721.5 | 1.53 |
+| RSA-2048 sign/decrypt | 759.4 | 612.0 | 659.4 | 1.08 |
+| RSA-2048 verify/encrypt | 33.56 | 30.21 | 18.86 | 0.62 |
 
 ### AArch64
 
@@ -128,43 +128,43 @@
 
 Throughput in MB/s, **higher is better**:
 
-| | crypton 1.1.5 | crypton 2.1.4 | OpenSSL | 2.1.4 / OpenSSL |
+| | crypton 1.1.5 | crypton 2.1.5 | OpenSSL | 2.1.5 / OpenSSL |
 | --- | ---: | ---: | ---: | ---: |
-| AES-128-GCM | 131 | 9487 | 11111 | 0.85 |
-| AES-256-GCM | 98 | 8149 | 9371 | 0.87 |
-| ChaCha20-Poly1305 | 786 | 2321 | 2304 | 1.01 |
-| SHA-1 | 1255 | 3382 | 3366 | 1.00 |
-| SHA-256 | 483 | 3396 | 3366 | 1.01 |
-| SHA-512 | 735 | 1878 | 1888 | 0.99 |
-| SHA3-256 | 557 | 1105 | 1100 | 1.00 |
+| AES-128-GCM | 127 | **12422** | 10846 | 1.15 |
+| AES-256-GCM | 98 | **9721** | 9197 | 1.06 |
+| ChaCha20-Poly1305 | 771 | 2319 | 2250 | 1.03 |
+| SHA-1 | 1209 | 3389 | 3361 | 1.01 |
+| SHA-256 | 474 | 3400 | 3362 | 1.01 |
+| SHA-512 | 730 | 1880 | 1883 | 1.00 |
+| SHA3-256 | 550 | 1075 | 1065 | 1.01 |
 
 Time per operation in microseconds, **lower is better**:
 
-| | crypton 1.1.5 | crypton 2.1.4 | OpenSSL | OpenSSL / 2.1.4 |
+| | crypton 1.1.5 | crypton 2.1.5 | OpenSSL | OpenSSL / 2.1.5 |
 | --- | ---: | ---: | ---: | ---: |
-| X25519 | 17.45 | **11.12** | 15.35 | 1.38 |
-| ECDH P-256 | 67.18 | **19.58** | 24.53 | 1.25 |
-| ECDH P-384 | 3192 | **72.35** | 373.5 | 5.16 |
-| Ed25519 sign | 13.13 | **7.60** | 13.26 | 1.75 |
-| Ed25519 verify | 18.02 | 17.93 | 34.97 | 1.95 |
-| ECDSA P-256 sign | 32.08 | **6.88** | 11.02 | 1.60 |
-| ECDSA P-256 verify | 96.07 | **27.42** | 32.77 | 1.20 |
-| ECDSA P-384 sign | 3253 | **124.2** | 396.8 | 3.19 |
-| ECDSA P-384 verify | 3912 | **203.7** | 333.3 | 1.64 |
-| RSA-2048 sign/decrypt | 452.0 | 465.3 | 321.3 | 0.69 |
-| RSA-2048 verify/encrypt | 18.23 | 15.27 | 8.44 | 0.55 |
+| X25519 | 18.27 | **12.22** | 15.53 | 1.27 |
+| ECDH P-256 | 68.70 | **20.43** | 24.77 | 1.21 |
+| ECDH P-384 | 3328 | **73.50** | 372.6 | 5.07 |
+| Ed25519 sign | 13.58 | **7.75** | 13.23 | 1.71 |
+| Ed25519 verify | 18.28 | 18.17 | 34.76 | 1.91 |
+| ECDSA P-256 sign | 31.97 | **6.55** | 10.92 | 1.67 |
+| ECDSA P-256 verify | 95.63 | **26.80** | 32.68 | 1.22 |
+| ECDSA P-384 sign | 3219 | **124.1** | 394.2 | 3.18 |
+| ECDSA P-384 verify | 3870 | **203.1** | 326.5 | 1.61 |
+| RSA-2048 sign/decrypt | 447.9 | 460.1 | 319.9 | 0.70 |
+| RSA-2048 verify/encrypt | 18.23 | 15.12 | 8.405 | 0.56 |
 
 ### What the numbers say
 
-There are two changes behind the 1.1.5 column and the 2.1.4 one, not a
+There are two changes behind the 1.1.5 column and the 2.1.5 one, not a
 single steady improvement.
 
 The first, in 2.0.0, was a rewrite: the bulk algorithms moved into C, the
 curves other than P-256 moved out of Haskell `Integer` arithmetic, and
 everything that touches a secret was made to take the same time whatever the
 secret is.  1.1.5 had no AArch64 code of its own at all, which is why AES-GCM
-there is seventy times what it was, and on x86-64 it had AES-NI and nothing
-else.
+there is close to a hundred times what it was, and on x86-64 it had AES-NI
+and nothing else.
 
 The second, from 2.1.0 onwards, is assembly, for the operations where C
 cannot reach.  Which of the two a row owes its gain to is not the same
@@ -184,15 +184,17 @@
 it is Apache-2.0 only, and Intel and CloudFlare hold copyright in it besides
 OpenSSL, so nobody is in a position to relicense it.
 
-Where crypton is behind, which is now one row on one architecture and the
-AES-GCM rows on the other, it is behind for two reasons.
+Where crypton is behind, which is now the RSA rows on both architectures and
+ECDSA P-256 verification on x86-64, there is one reason.  The AES-GCM rows
+were the other half of this section until 2.1.5; they are ahead on both
+machines now, and what the instructions do is still worth setting out.
 
 *RSA.*  2.0.0 made signing slower than 1.1.5 on purpose: its modular
 exponentiation stopped indexing a table with the bits of the exponent, and
 hiding the exponent is what the difference bought.  On x86-64 that cost is
 more than repaid -- s2n-bignum's Montgomery multiplication is twice the C's,
 because the C cannot form the two carry chains `ADCX` and `ADOX` give, and
-2.1.4 signs in less than 1.1.5 took while keeping what 2.0.0 gained.  On
+2.1.5 signs in less than 1.1.5 took while keeping what 2.0.0 gained.  On
 AArch64 there is nothing to use: s2n-bignum has no generic routine for it,
 and the same five that help on x86-64 measure level with the C there, so the
 C stays and the gap with it.  No portable C closes that gap either -- the
@@ -209,7 +211,7 @@
 BoringSSL and AWS-LC is Apache-2.0 and s2n-bignum has no GCM, so both files are
 crypton's own.
 
-Having the 256-bit one is where the 1.48 in the x86-64 table comes from, and
+Having the 256-bit one is where the 1.49 in the x86-64 table comes from, and
 it is narrower than it sounds.  The EPYC 7763 is Zen 3: VAES and VPCLMULQDQ,
 no AVX-512.  OpenSSL's x86-64 AES-GCM is `aesni-gcm-x86_64.pl`, which is
 128-bit -- its `vaesenc`s are the VEX encoding of `AESENC` on `xmm`, and
@@ -218,7 +220,7 @@
 takes a block at a time where crypton takes two.  The same idea as theirs,
 one step further down the feature ladder; not a better one.
 
-The 512-bit path arrived after 2.1.2, so it is in the 2.1.4 column -- but
+The 512-bit path arrived after 2.1.2, so it is in the 2.1.5 column -- but
 neither machine in the tables above has AVX-512, so neither column shows it.
 On the runners that do, measured over 16 KiB in MB/s: an EPYC 9V45 (Zen 5)
 goes from 9616 to 14268 with it, a Xeon 6973P-C from 8095 to 9848, a Xeon
@@ -228,15 +230,22 @@
 instructions are two passes through a 256-bit datapath, so the wider encoding
 buys nothing there and costs a little.
 
-AArch64 has no counterpart to any of these, which is where the 0.85 on its
-AES-GCM rows comes from -- and, the other way about, why the x86-64 rows are
-at 1.48 and 1.45.
+AArch64 has no counterpart to any of these: one AES block and one GHASH
+multiplication at a time is all the instruction set offers.  Its AES-GCM
+rows were 0.85 and 0.87 until 2.1.5, for that reason.  What closed it was
+not width but the GHASH's representation -- H is twisted once at key setup
+so that GCM's bit reflection is already undone, which turns a reduction of
+some twenty-five shifts and XORs into two PMULL and six EOR and makes
+Karatsuba worth taking.  The scheme is ARM's, from the BSD-3-Clause part of
+[AArch64cryptolib](https://github.com/ARM-software/AArch64cryptolib),
+written out in crypton's own intrinsics.  The AES there is ahead of
+OpenSSL's and always was; it was the GHASH beside it that was behind.
 
 One row wants a word of its own: crypton's `Ed25519.sign` derives the public
 key from the secret key every time it signs, so that a caller who passes a
 public key that does not match cannot be made to leak the private one.  That
 costs a second scalar multiplication, which OpenSSL's signing does not pay --
-and the row is still 1.75 on AArch64 and 1.80 on x86-64, so the safety is had
+and the row is still 1.71 on AArch64 and 1.81 on x86-64, so the safety is had
 for nothing here rather than paid for.
 
 SHA-1 is in the tables because a number of protocols and file formats still
diff --git a/cbits/aes/LICENSE.fusion b/cbits/aes/LICENSE.fusion
new file mode 100644
--- /dev/null
+++ b/cbits/aes/LICENSE.fusion
@@ -0,0 +1,29 @@
+Parts of cbits/aes/gcm_fused_x86.c follow the AES-GCM implementation in
+picotls, lib/fusion.c, which is under the MIT license reproduced below.
+The design is described by its author at
+
+    http://blog.kazuhooku.com/2020/06/quicaes-gcm-12.html
+    http://blog.kazuhooku.com/2020/06/quicaes-gcm-22.html
+
+and the source is at https://github.com/h2o/picotls.
+
+
+Copyright (c) 2020-2022 Fastly, Kazuho Oku
+
+Permission is hereby granted, free of charge, to any person obtaining a copy
+of this software and associated documentation files (the "Software"), to
+deal in the Software without restriction, including without limitation the
+rights to use, copy, modify, merge, publish, distribute, sublicense, and/or
+sell copies of the Software, and to permit persons to whom the Software is
+furnished to do so, subject to the following conditions:
+
+The above copyright notice and this permission notice shall be included in
+all copies or substantial portions of the Software.
+
+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
+AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
+LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING
+FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS
+IN THE SOFTWARE.
diff --git a/cbits/aes/gcm_fused_x86.c b/cbits/aes/gcm_fused_x86.c
--- a/cbits/aes/gcm_fused_x86.c
+++ b/cbits/aes/gcm_fused_x86.c
@@ -1,10 +1,18 @@
 /*
- * A fused AES-GCM for x86-64, written to the design Kazuho Oku sets out in
- * "QUICむけにAES-GCM実装を最適化した話": keep AES-NI issuing every clock and
- * fit everything else -- the additional data, the tag, the QUIC header
- * protection mask -- into the gaps it leaves.  Written in C with intrinsics
- * rather than assembly, for the same reason he gives: the scheduling is what
- * is complicated here, and it has to stay readable to stay correct.
+ * A fused AES-GCM for x86-64, following the design Kazuho Oku sets out in
+ * "QUICむけにAES-GCM実装を最適化した話" and implements in picotls's
+ * lib/fusion.c: keep AES-NI issuing every clock and fit everything else --
+ * the additional data, the tag, the QUIC header protection mask -- into the
+ * gaps it leaves.  Written in C with intrinsics rather than assembly, for
+ * the same reason he gives: the scheduling is what is complicated here, and
+ * it has to stay readable to stay correct.
+ *
+ * Parts of this file follow fusion closely enough to say so: `loadn` and the
+ * two tables it reads, `loadn_page_end` and `storen` are its `loadn128`,
+ * `loadn_end_of_page` and `storen128` in another spelling, and the
+ * reduction of a 256-bit product is the sequence fusion takes from Gueron's
+ * "AES-GCM for Efficient Authenticated Encryption".  fusion is under the MIT
+ * license, which is in cbits/aes/LICENSE.fusion beside this file.
  *
  * The powers of H are built once per key, so the additional data, the
  * ciphertext and the length block are absorbed against them in batches that
diff --git a/cbits/aes/x86ni.c b/cbits/aes/x86ni.c
--- a/cbits/aes/x86ni.c
+++ b/cbits/aes/x86ni.c
@@ -209,13 +209,23 @@
 	return v;
 }
 
+/* memcpy rather than a cast, as everything else that moves bytes between a
+ * crypton structure and a word does since block128 was packed.  The cast
+ * this replaces was written in 2014, when block128 was a plain union and
+ * taking a __m128i * to one promised nothing the type did not already
+ * offer.  Packing it dropped its alignment to one, and the promise with it:
+ * gcc has reported the cast ever since, and it is right to -- the attribute
+ * on the local below is what makes the promise true, and nothing obliges
+ * the next edit to keep it.  Sixteen bytes of memcpy between a __m128i and
+ * a sixteen-byte object is one movdqu, or nothing at all when both stay in
+ * registers. */
 TARGET_AESNI
 static __m128i gfmul_generic(__m128i tag, const table_4bit htable)
 {
-	aes_block _t ALIGNMENT(16);
-	_mm_store_si128((__m128i *) &_t, tag);
+	aes_block _t;
+	memcpy(&_t, &tag, sizeof _t);
 	crypton_aes_generic_gf_mul(&_t, htable);
-	tag = _mm_load_si128((__m128i *) &_t);
+	memcpy(&tag, &_t, sizeof tag);
 	return tag;
 }
 
diff --git a/cbits/crypton_pbkdf2.c b/cbits/crypton_pbkdf2.c
--- a/cbits/crypton_pbkdf2.c
+++ b/cbits/crypton_pbkdf2.c
@@ -46,6 +46,7 @@
 
 /* Internal function/type names for hash-specific things. */
 #define HMAC_CTX(_name) HMAC_ ## _name ## _ctx
+#define DIGEST_FITS(_name) pbkdf2_ ## _name ## _digest_fits_in_a_block
 #define HMAC_INIT(_name) HMAC_ ## _name ## _init
 #define HMAC_UPDATE(_name) HMAC_ ## _name ## _update
 #define HMAC_FINAL(_name) HMAC_ ## _name ## _final
@@ -79,6 +80,16 @@
  */
 #define DECL_PBKDF2(_name, _blocksz, _hashsz, _ctx,                           \
                     _init, _update, _xform, _final, _xcpy, _xtract, _xxor)    \
+  /* HMAC_INIT below shortens a key longer than the block by hashing it,     \
+   * which writes _hashsz bytes into a buffer of _blocksz.  An instantiation \
+   * whose digest is larger than its block would overflow that buffer, and   \
+   * would do it before any check inside the function could say so -- which  \
+   * is where the check used to be.  Refuse such an instantiation here       \
+   * instead, in front of the person writing it.  The three below are        \
+   * SHA-1, SHA-256 and SHA-512, whose digests are 20, 32 and 64 bytes       \
+   * against blocks of 64, 64 and 128. */                                    \
+  typedef char DIGEST_FITS(_name)[(_hashsz) <= (_blocksz) ? 1 : -1];          \
+                                                                              \
   typedef struct {                                                            \
     _ctx inner;                                                               \
     _ctx outer;                                                               \
@@ -101,9 +112,6 @@
       nkey = _hashsz;                                                         \
     }                                                                         \
                                                                               \
-    /* Standard doesn't cover case where blocksz < hashsz. */                 \
-    assert(nkey <= _blocksz);                                                 \
-                                                                              \
     /* Right zero-pad short keys. */                                          \
     if (k != key)                                                             \
       memcpy(k, key, nkey);                                                   \
@@ -192,7 +200,16 @@
                      uint8_t *out, size_t nout)                               \
   {                                                                           \
     assert(iterations);                                                       \
-    assert(out && nout);                                                      \
+    assert(out);                                                              \
+                                                                              \
+    /* Zero bytes of derived key is zero bytes of work.  RFC 8018 asks for a  \
+     * positive dkLen and the loop below would write a block regardless, so   \
+     * this used to be `assert(out && nout)` -- which aborts the process, in  \
+     * a library built without NDEBUG, on a length the caller chose.          \
+     * Crypto.KDF.PBKDF2's own tryGenerate returns an empty result for this,  \
+     * so return and let the two agree. */                                    \
+    if (nout == 0)                                                            \
+      return;                                                                 \
                                                                               \
     /* Starting point for inner loop. */                                      \
     HMAC_CTX(_name) ctx;                                                      \
diff --git a/cbits/crypton_sha256.h b/cbits/crypton_sha256.h
--- a/cbits/crypton_sha256.h
+++ b/cbits/crypton_sha256.h
@@ -51,6 +51,18 @@
 
 void crypton_sha256_init(struct sha256_ctx *ctx);
 void crypton_sha256_update(struct sha256_ctx *ctx, const uint8_t *data, uint32_t len);
+/* The pointers are all required to be non-null, which is said here so that
+ * the compiler knows it too.  Both of these write their digest through a
+ * loop -- store_be32(out + 4 * i, ...) -- where sha1 and md5 write theirs at
+ * constant offsets, and that is the difference that makes gcc's
+ * -Wstringop-overflow reason about out being null: with
+ * -fsanitize=undefined, UndefinedBehaviorSanitizer inserts a null check
+ * before memcpy, because glibc declares memcpy nonnull, and the check puts a
+ * null path in front of the warning pass, which then reports writing into
+ * "a region of size 0" at "address zero".  Saying the pointer is never null
+ * removes the path rather than the warning.  It costs nothing: compiled as
+ * the package compiles it, the assembly is identical with and without. */
+__attribute__((nonnull))
 void crypton_sha256_finalize(struct sha256_ctx *ctx, uint8_t *out);
 void crypton_sha256_finalize_prefix(struct sha256_ctx *ctx, const uint8_t *data, uint32_t len, uint32_t n, uint8_t *out);
 
diff --git a/cbits/crypton_sha512.h b/cbits/crypton_sha512.h
--- a/cbits/crypton_sha512.h
+++ b/cbits/crypton_sha512.h
@@ -50,6 +50,18 @@
 
 void crypton_sha512_init(struct sha512_ctx *ctx);
 void crypton_sha512_update(struct sha512_ctx *ctx, const uint8_t *data, uint32_t len);
+/* The pointers are all required to be non-null, which is said here so that
+ * the compiler knows it too.  Both of these write their digest through a
+ * loop -- store_be32(out + 4 * i, ...) -- where sha1 and md5 write theirs at
+ * constant offsets, and that is the difference that makes gcc's
+ * -Wstringop-overflow reason about out being null: with
+ * -fsanitize=undefined, UndefinedBehaviorSanitizer inserts a null check
+ * before memcpy, because glibc declares memcpy nonnull, and the check puts a
+ * null path in front of the warning pass, which then reports writing into
+ * "a region of size 0" at "address zero".  Saying the pointer is never null
+ * removes the path rather than the warning.  It costs nothing: compiled as
+ * the package compiles it, the assembly is identical with and without. */
+__attribute__((nonnull))
 void crypton_sha512_finalize(struct sha512_ctx *ctx, uint8_t *out);
 void crypton_sha512_finalize_prefix(struct sha512_ctx *ctx, const uint8_t *data, uint32_t len, uint32_t n, uint8_t *out);
 
diff --git a/cbits/p256/p256.c b/cbits/p256/p256.c
--- a/cbits/p256/p256.c
+++ b/cbits/p256/p256.c
@@ -329,6 +329,22 @@
   crypton_p256_int U = *MOD;
   crypton_p256_int V = *a;
 
+  /* Zero has no inverse, and the loop below never finds that out: V stays
+     even forever, so it is halved forever, and the only break is in the
+     branch both U and V have to be odd to reach.  The other input without an
+     inverse is MOD itself -- 2*MOD does not fit in 256 bits, so there is no
+     third -- and that one already leaves here as zero, which is also what
+     crypton_p256_modinv's constant-time counterpart returns.  Answer the same
+     for zero rather than not answering.
+
+     Reachable: Crypto.PubKey.ECC.P256 exports scalarInv, and scalarFromBinary
+     accepts any 256 bits.  A hang inside a foreign call cannot be interrupted
+     by System.Timeout either. */
+  if (crypton_p256_is_zero(a)) {
+    crypton_p256_clear(b);
+    return;
+  }
+
   for (;;) {
     if (crypton_p256_is_even(&U)) {
       crypton_p256_shr1(&U, 0, &U);
diff --git a/crypton.cabal b/crypton.cabal
--- a/crypton.cabal
+++ b/crypton.cabal
@@ -1,9 +1,24 @@
 cabal-version:      3.0
 name:               crypton
-version:            2.1.5
-license:            BSD-3-Clause
-license-file:       LICENSE
-copyright:          Vincent Hanquez <vincent@snarc.org>
+version:            2.1.6
+-- crypton's own code is BSD-3-Clause.  The parts of
+-- cbits/aes/gcm_fused_x86.c that follow picotls's fusion are MIT, and the
+-- vendored s2n-bignum assembly in cbits/s2n is taken under ISC; each has
+-- its licence beside it, and they are listed below.  The CRYPTOGAMS
+-- assembly in cbits/asm is taken under its BSD-3-Clause option, and the
+-- AArch64 multiply-accumulate loop in cbits/crypton_bignum.h follows Go's,
+-- which is BSD-3-Clause too; the first term already covers both.
+license:            BSD-3-Clause AND MIT AND ISC
+license-files:
+    LICENSE
+    cbits/LICENSE.go
+    cbits/aes/LICENSE.fusion
+    cbits/asm/LICENSE.cryptogams
+    cbits/s2n/LICENSE
+copyright:
+    2006-2022 Vincent Hanquez <vincent@snarc.org> and contributors,
+    2023-2026 Kazu Yamamoto <kazu@iij.ad.jp>
+
 maintainer:         Kazu Yamamoto <kazu@iij.ad.jp>
 author:             Vincent Hanquez <vincent@snarc.org>
 stability:          experimental
@@ -26,8 +41,6 @@
     cbits/curve25519/*.h
     cbits/aes/armv8_impl.c
     cbits/aes/x86ni_impl.c
-    cbits/LICENSE.go
-    cbits/asm/LICENSE.cryptogams
     cbits/asm/README.md
     cbits/asm/aesni-gcm-x86_64.pl
     cbits/asm/arm-xlate.pl
@@ -62,7 +75,6 @@
     cbits/include32/p256/*.h
     cbits/include64/p256/*.h
     cbits/s2n/COMMIT
-    cbits/s2n/LICENSE
     cbits/s2n/README.md
     cbits/s2n/arm/*.S
     cbits/p256/*.h
diff --git a/tests/KDF/PBKDF2Spec.hs b/tests/KDF/PBKDF2Spec.hs
--- a/tests/KDF/PBKDF2Spec.hs
+++ b/tests/KDF/PBKDF2Spec.hs
@@ -110,7 +110,18 @@
                 `shouldBe` refused
             PBKDF2.tryFastPBKDF2_SHA512 (PBKDF2.Parameters 1 (-1)) badPass badSalt
                 `shouldBe` refused
+        -- Zero is not rejected: asking for no key is asking for no work, and
+        -- that is what the slow path has always answered.  The fast ones go
+        -- straight to C, where `assert(out && nout)` took the process down
+        -- on a length the caller chose -- a library built without NDEBUG
+        -- keeps its assertions.  All four agree now.
+        it "derives nothing when asked for nothing" $ do
+            slow none `shouldBe` ""
+            fast1 none `shouldBe` ""
+            fast256 none `shouldBe` ""
+            fast512 none `shouldBe` ""
   where
+    none = PBKDF2.Parameters 1 0
     badPrf = PBKDF2.prfHMAC SHA256
     badPass = "password" :: ByteString
     badSalt = "salt" :: ByteString
diff --git a/tests/PubKey/P256Spec.hs b/tests/PubKey/P256Spec.hs
--- a/tests/PubKey/P256Spec.hs
+++ b/tests/PubKey/P256Spec.hs
@@ -157,6 +157,21 @@
                     [ eqTest "scalarZero" P256.scalarZero inv0
                     , eqTest "scalarN" P256.scalarZero invN
                     ]
+        -- The same two for the variable-time inverse, which is exported and
+        -- which scalarFromBinary will happily hand a zero.  It used not to
+        -- return at all on that: in the binary extended Euclid below it, zero
+        -- stays even and is halved forever, and the loop's only exit is in
+        -- the branch both operands must be odd to reach.  Not even
+        -- System.Timeout gets a program out of that, the hang being inside a
+        -- foreign call.  The properties above step around it with a
+        -- precondition; this one walks into it.
+        prop "inv-zero" $
+            let inv0 = P256.scalarInv P256.scalarZero
+                invN = P256.scalarInv P256.scalarN
+             in propertyHold
+                    [ eqTest "scalarZero" P256.scalarZero inv0
+                    , eqTest "scalarN" P256.scalarZero invN
+                    ]
     describe "point" $ do
         prop "marshalling" $ \rx ry ->
             let p = P256.pointFromIntegers (unP256 rx, unP256 ry)
