[TUHS] Another awk version is now available
Arnold Robbins via TUHS
tuhs at tuhs.org
Sun Sep 27 17:05:55 AEST 2026
Steffen Nurpmeso via TUHS <tuhs at tuhs.org> wrote:
> Constrained by that neither any ISO C nor (thus) any POSIX offer
> the possibility to truly dig into "multibyte character sets" for
> eq regular expressions the necessary way.
> (Collation, equivalence classes.
True. The charset library used in gawk supports equivalence classes
using pregenerated tables for Unicode; the lack of standard APIs for
collating sequences and equivalence classes is a problem. But not
a new one.
> Let alone the full power of
> U(nicode)T(echnical)S(tandard)18 Unicode regular expressions.
It's been a while since I looked. The unicode.org web site is (or was)
a maze of twisty litle passages, all alike, but ISTR they have a regex
library that fully implements their standard. However, I don't know if
it supports the POSIX BRE/ERE flavors and POSIX leftmost-longest rules.
> And also constrained by "multibyte character set" meaning NUL
> terminated string compatible i would think, given that you
> signalled "won't fix" for nawk giving this argument as the reason
I'm not sure what you're referring to here. The One True Awk (the
direct descendant of the original Unix awk) uses C strings
internally, so doing anything with NUL is a problem. "Fixing" that
is too much work, and I no longer maintain that code, somebody else does.
> Thanks for decades of awk work, i use awk very frequently.
You are welcome.
Arnold
More information about the TUHS
mailing list