Liking cljdoc? Tell your friends :D

datahike.pg.datetime.tokens

PostgreSQL's datetime token tables, transcribed from the pinned 17.7.

TWO tables, and they must stay two. datetktbl (datetime.c:105-179) is compiled into the backend and holds months, weekdays, am/pm, ad/bc, the RESERV words (now, today, epoch, infinity, allballs) and the ISO unit letters. The ZONE ABBREVIATIONS are not in it -- datetime.c:100-103 says so outright: "The static table contains no TZ, DTZ, or DYNTZ entries; rather those are loaded from configuration files" -- they come from src/timezone/tznames/Default, and DecodeTimezoneAbbrev is consulted BEFORE DecodeSpecial (datetime.c:1304-1309).

That order is not negotiable, and the comment at datetime.c:3203-3208 says why: tzdb deliberately contains zone NAMES identical to offset ABBREVIATIONS. With the Default set the two token sets happen not to collide at all (checked below), so the precedence is currently a no-op -- it stops being one the moment timezone_abbreviations is settable, which is why it is written down rather than relied upon.

An abbreviation is a FIXED OFFSET; a zone NAME is a RULE. pst here is -28800 seconds, always, with no DST -- while PST8PDT is a tzdb LINK to America/Los_Angeles and does observe DST. Resolving an abbreviation through ZoneId/of with SHORT_IDS (which maps PST to America/Los_Angeles) gets every July timestamp wrong by an hour, with no error, which is what this table exists to stop.

Keys are LOWERCASE. ParseDateTime lowercases every alpha run with pg_tolower, which is ASCII-only (pgstrcasecmp.c:122-129), and tzparser.c:79-83 lowercases every abbreviation with the comment "must match datetime.c's conversion". So a lookup must fold ASCII only -- never clojure.string/lower-case, which is locale-sensitive and breaks under a Turkish default locale.

GENERATED -- the extraction commands are in the comment below.

PostgreSQL's datetime token tables, transcribed from the pinned 17.7.

TWO tables, and they must stay two. `datetktbl` (datetime.c:105-179)
is compiled into the backend and holds months, weekdays, am/pm, ad/bc,
the RESERV words (`now`, `today`, `epoch`, `infinity`, `allballs`) and
the ISO unit letters. The ZONE ABBREVIATIONS are not in it --
datetime.c:100-103 says so outright: "The static table contains no TZ,
DTZ, or DYNTZ entries; rather those are loaded from configuration
files" -- they come from `src/timezone/tznames/Default`, and
`DecodeTimezoneAbbrev` is consulted BEFORE `DecodeSpecial`
(datetime.c:1304-1309).

That order is not negotiable, and the comment at datetime.c:3203-3208
says why: tzdb deliberately contains zone NAMES identical to offset
ABBREVIATIONS. With the `Default` set the two token sets happen not to
collide at all (checked below), so the precedence is currently a
no-op -- it stops being one the moment `timezone_abbreviations` is
settable, which is why it is written down rather than relied upon.

An abbreviation is a FIXED OFFSET; a zone NAME is a RULE. `pst` here
is -28800 seconds, always, with no DST -- while `PST8PDT` is a tzdb
LINK to America/Los_Angeles and does observe DST. Resolving an
abbreviation through `ZoneId/of` with `SHORT_IDS` (which maps PST to
America/Los_Angeles) gets every July timestamp wrong by an hour, with
no error, which is what this table exists to stop.

Keys are LOWERCASE. `ParseDateTime` lowercases every alpha run with
`pg_tolower`, which is ASCII-only (pgstrcasecmp.c:122-129), and
`tzparser.c:79-83` lowercases every abbreviation with the comment
"must match datetime.c's conversion". So a lookup must fold ASCII
only -- never `clojure.string/lower-case`, which is locale-sensitive
and breaks under a Turkish default locale.

GENERATED -- the extraction commands are in the comment below.
raw docstring

ascii-lowerclj

(ascii-lower s)

pg_tolower's fold: A-Z only. clojure.string/lower-case is locale-sensitive -- under a Turkish default locale it maps I to a dotless i and every abbreviation starting with I stops resolving.

`pg_tolower`'s fold: A-Z only. `clojure.string/lower-case` is
locale-sensitive -- under a Turkish default locale it maps `I` to a
dotless i and every abbreviation starting with I stops resolving.
sourceraw docstring

datetktblclj

datetktbl (datetime.c:105-179), 72 entries: {token -> [type value]}.

The type and value are kept as KEYWORDS rather than resolved to the C integers. The numbers only exist because C needed bitmask arithmetic over them; what the decoder actually wants is a tag and a set, and a keyword says which rule applies without a lookup table of its own. A numeric value (a month number, a weekday number, an offset in seconds) stays a number.

`datetktbl` (datetime.c:105-179), 72 entries: {token -> [type value]}.

The `type` and `value` are kept as KEYWORDS rather than resolved to
the C integers. The numbers only exist because C needed bitmask
arithmetic over them; what the decoder actually wants is a tag and a
set, and a keyword says which rule applies without a lookup table of
its own. A numeric `value` (a month number, a weekday number, an
offset in seconds) stays a number.
sourceraw docstring

prefix-safe?clj

(prefix-safe? table)

Is a plain map lookup equivalent to datebsearch's 10-character prefix match for this table? True when no key reaches tokmaxlen.

Is a plain map lookup equivalent to `datebsearch`'s 10-character
prefix match for this table? True when no key reaches `tokmaxlen`.
sourceraw docstring

special-tokenclj

(special-token s)

DecodeSpecial (datetime.c:3148-3188): look s up in datetktbl. Returns [type value] or nil. s must already be ASCII-lowercased, as ParseDateTime leaves it.

`DecodeSpecial` (datetime.c:3148-3188): look `s` up in `datetktbl`.
Returns `[type value]` or nil. `s` must already be ASCII-lowercased,
as `ParseDateTime` leaves it.
sourceraw docstring

tokmaxlenclj

source

zone-abbrevclj

(zone-abbrev s)

DecodeTimezoneAbbrev (datetime.c:3091-3146): look s up in the abbreviation table ONLY -- never in datetktbl, and never through pg_tzset. Returns the entry or nil; nil is UNKNOWN_FIELD, which is what sends the caller on to DecodeSpecial and then to pg_tzset.

`DecodeTimezoneAbbrev` (datetime.c:3091-3146): look `s` up in the
abbreviation table ONLY -- never in `datetktbl`, and never through
`pg_tzset`. Returns the entry or nil; nil is `UNKNOWN_FIELD`, which is
what sends the caller on to `DecodeSpecial` and then to `pg_tzset`.
sourceraw docstring

zone-abbrevsclj

src/timezone/tznames/Default, 195 entries: 97 TZ, 48 DTZ, 50 DYNTZ.

TZ a fixed offset, standard time: {:type :tz :offset <seconds east>} DTZ a fixed offset, daylight time; sets tm_isdst (datetime.c:1406-1418) DYNTZ resolved against a zone at the instant in question (DetermineTimeZoneAbbrevOffset): {:type :dyntz :zone "..."}

ConvertTimeZoneAbbrevs (datetime.c:4873-4951) assigns DYNTZ when the second column is a zone name and is_dst ? DTZ : TZ otherwise.

The offsets here are SECONDS EAST, as the file writes them. PostgreSQL stores tzp internally as seconds WEST and negates on the way in (*tzp = -tz, datetime.c:3083) -- getting that backwards is a silent double-offset error, so the sign convention is stated at every boundary.

`src/timezone/tznames/Default`, 195 entries: 97 TZ, 48 DTZ, 50 DYNTZ.

TZ   a fixed offset, standard time: {:type :tz :offset <seconds east>}
DTZ  a fixed offset, daylight time; sets tm_isdst
     (datetime.c:1406-1418)
DYNTZ  resolved against a zone at the instant in question
     (`DetermineTimeZoneAbbrevOffset`): {:type :dyntz :zone "..."}

`ConvertTimeZoneAbbrevs` (datetime.c:4873-4951) assigns DYNTZ when the
second column is a zone name and `is_dst ? DTZ : TZ` otherwise.

The offsets here are SECONDS EAST, as the file writes them. PostgreSQL
stores `tzp` internally as seconds WEST and negates on the way in
(`*tzp = -tz`, datetime.c:3083) -- getting that backwards is a silent
double-offset error, so the sign convention is stated at every
boundary.
sourceraw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close