Liking cljdoc? Tell your friends :D

datahike.pg.interval.core

The interval CARRIER, and interval_out.

A deftype, deliberately, and not a defrecord. The consolidation plan warns that "a record's structural =/hash would split GROUP BY and joins", and that is true -- but it is a reason to choose the carrier's equality, not a reason to avoid a carrier. PostgreSQL compares '1 mon' and '30 days' as EQUAL (interval_cmp_value: a month is 30 days, a day is 24 hours) while PRINTING them differently, so:

structural equality splits two values PostgreSQL groups; canonical TEXT cannot be the key either, for the same reason.

clojure.core/= routes a non-collection, non-Number object to .equals, and clojure.core/hash routes a non-IHashEq, non-Number, non-String object to .hashCode. So implementing those two over interval_cmp_value makes the EXISTING machinery correct, including the parts no comparator can reach: GROUP BY is datahike's own group-by over the non-aggregate :find elements, and DISTINCT, clojure.set/intersection, contains? on a set and datahike's hash join are all clojure.core calls. None of them takes a comparator; all of them are served by .hashCode.

Comparable covers ORDER BY, the comparison operators, BETWEEN and min/max, through fns/order-cmp's existing :else (compare a b) -- so that function needs no interval branch at all. The same route bits.clj's PgBit takes.

THE SPAN NEEDS MORE THAN 64 BITS. month * 30 + day can reach 6.4e10 days, and multiplying that by 86,400,000,000 µs overflows int64 -- which is why the C uses INT128. compareTo therefore works in BigInteger. hashCode does NOT: interval_hash deliberately narrows the INT128 to its low 64 bits (int128_to_int64) before hashing, so two spans differing by exactly 2^64 µs hash together in PostgreSQL, and matching that is what keeps our hashing consistent with a real server's.

toString is NOT the SQL rendering. Output goes through types/->pg-text, which is the single funnel for ::text, array elements, record fields, COPY, pg_dump and to_jsonb.

The interval CARRIER, and `interval_out`.

A `deftype`, deliberately, and not a `defrecord`. The consolidation
plan warns that "a record's structural `=`/hash would split GROUP BY
and joins", and that is true -- but it is a reason to choose the
carrier's equality, not a reason to avoid a carrier. PostgreSQL
compares `'1 mon'` and `'30 days'` as EQUAL (`interval_cmp_value`:
a month is 30 days, a day is 24 hours) while PRINTING them
differently, so:

  structural equality splits two values PostgreSQL groups;
  canonical TEXT cannot be the key either, for the same reason.

`clojure.core/=` routes a non-collection, non-Number object to
`.equals`, and `clojure.core/hash` routes a non-`IHashEq`,
non-Number, non-String object to `.hashCode`. So implementing those
two over `interval_cmp_value` makes the EXISTING machinery correct,
including the parts no comparator can reach: `GROUP BY` is
datahike's own `group-by` over the non-aggregate `:find` elements,
and `DISTINCT`, `clojure.set/intersection`, `contains?` on a set and
datahike's hash join are all `clojure.core` calls. None of them
takes a comparator; all of them are served by `.hashCode`.

`Comparable` covers ORDER BY, the comparison operators, BETWEEN and
min/max, through `fns/order-cmp`'s existing `:else (compare a b)` --
so that function needs no interval branch at all. The same route
`bits.clj`'s `PgBit` takes.

THE SPAN NEEDS MORE THAN 64 BITS. `month * 30 + day` can reach
6.4e10 days, and multiplying that by 86,400,000,000 µs overflows
int64 -- which is why the C uses INT128. `compareTo` therefore works
in `BigInteger`. `hashCode` does NOT: `interval_hash` deliberately
narrows the INT128 to its low 64 bits (`int128_to_int64`) before
hashing, so two spans differing by exactly 2^64 µs hash together in
PostgreSQL, and matching that is what keeps our hashing consistent
with a real server's.

`toString` is NOT the SQL rendering. Output goes through
`types/->pg-text`, which is the single funnel for `::text`, array
elements, record fields, COPY, pg_dump and `to_jsonb`.
raw docstring

datahike.pg.interval.decode

DecodeInterval (datetime.c:3364-3749) and the Adjust* helpers (560-670).

It reads the SAME lexed fields the datetime decoders read -- it is a fourth consumer of datetime/lex.clj, with buflen 256 rather than 153 or 129 -- and then interprets them by a different set of rules.

THE LOOP RUNS RIGHT TO LEFT (datetime.c:3397). Units FOLLOW their values in an interval literal, so '1 day' is read as day and then 1, with the unit setting type for the field about to be read. Nothing else in the datetime family works this way, and reversing it is not a detail: parsing_unit_val, AGO's last-field-only rule and DTK_HOUR's hand-off all depend on the direction.

The accumulator is pg_itm_in (timestamp.h:82): {:usec :mday :mon :year}, where YEAR is kept separate from MONTH during decoding and folded in only at the end by itmin2interval. Keeping them apart is what lets AdjustYears multiply by 10, 100 and 1000 for decade, century and millennium without losing the int32 range.

OVERFLOW IS CHECKED AT EVERY STEP in the C, with pg_add_s32_overflow and friends, and each failure is DTERR_FIELD_OVERFLOW -- which interval_in then REMAPS to DTERR_INTERVAL_OVERFLOW, SQLSTATE 22015, not 22008 (timestamp.c:932-941). Clojure longs do not overflow at int32, so the bounds are checked explicitly; a missing check is a silently wrapped value rather than an error.

`DecodeInterval` (datetime.c:3364-3749) and the `Adjust*` helpers
(560-670).

It reads the SAME lexed fields the datetime decoders read -- it is a
fourth consumer of `datetime/lex.clj`, with `buflen` 256 rather than
153 or 129 -- and then interprets them by a different set of rules.

THE LOOP RUNS RIGHT TO LEFT (datetime.c:3397). Units FOLLOW their
values in an interval literal, so `'1 day'` is read as `day` and
then `1`, with the unit setting `type` for the field about to be
read. Nothing else in the datetime family works this way, and
reversing it is not a detail: `parsing_unit_val`, AGO's
last-field-only rule and `DTK_HOUR`'s hand-off all depend on the
direction.

The accumulator is `pg_itm_in` (timestamp.h:82): `{:usec :mday :mon
:year}`, where YEAR is kept separate from MONTH during decoding and
folded in only at the end by `itmin2interval`. Keeping them apart is
what lets `AdjustYears` multiply by 10, 100 and 1000 for decade,
century and millennium without losing the int32 range.

OVERFLOW IS CHECKED AT EVERY STEP in the C, with
`pg_add_s32_overflow` and friends, and each failure is
DTERR_FIELD_OVERFLOW -- which `interval_in` then REMAPS to
DTERR_INTERVAL_OVERFLOW, SQLSTATE 22015, not 22008
(timestamp.c:932-941). Clojure longs do not overflow at int32, so
the bounds are checked explicitly; a missing check is a silently
wrapped value rather than an error.
raw docstring

datahike.pg.interval.tokens

deltatktbl (datetime.c:187-251), PostgreSQL's INTERVAL unit table.

A THIRD token table, beside datetime/tokens.clj's two. It is not a superset of either and the three disagree on purpose:

m is MINUTE here and MONTH in datetktbl mm is only in datetktbl j is only in datetktbl

And the disagreement is REACHABLE, because DecodeInterval's DTK_STRING arm tries DecodeUnits (this table) and then FALLS BACK to DecodeSpecial (datetktbl) -- datetime.c:3659-3663. So '1 m' is one minute and '1 mm' is also one minute, by two different routes. Three tables, consulted in a fixed order.

A PLAIN MAP LOOKUP IS WRONG HERE, unlike for the two datetime tables where datetime.tokens/prefix-safe? holds. datebsearch compares with strncmp(key, token, TOKMAXLEN) and TOKMAXLEN is 10, and FIVE keys in this table are exactly ten characters:

microsecon millennium millisecon timezone_h timezone_m

so each is a genuine PREFIX. Checked on the oracle: '1 millenniumXYZ'::interval is 1000 years and '2 microsecondsXX'::interval is 00:00:00.000002.

FOUR UNITS RESOLVE HERE AND ARE STILL ERRORS. DTK_QUARTER and the three timezone units have no arm in DecodeInterval's switch (datetime.c:3570-3650), so they reach its default and raise 22007. '1 qtr', '1 quarter', '1 timezone', '1 timezone_hour' and '1 timezone_minute' are all invalid intervals even though the lookup succeeds. A table-driven decoder that trusts this table would accept every one of them.

GENERATED -- the extraction command is in the comment below.

`deltatktbl` (datetime.c:187-251), PostgreSQL's INTERVAL unit table.

A THIRD token table, beside `datetime/tokens.clj`'s two. It is not a
superset of either and the three disagree on purpose:

  `m`   is MINUTE here and MONTH in `datetktbl`
  `mm`  is only in `datetktbl`
  `j`   is only in `datetktbl`

And the disagreement is REACHABLE, because `DecodeInterval`'s
DTK_STRING arm tries `DecodeUnits` (this table) and then FALLS BACK
to `DecodeSpecial` (`datetktbl`) -- datetime.c:3659-3663. So `'1 m'`
is one minute and `'1 mm'` is also one minute, by two different
routes. Three tables, consulted in a fixed order.

A PLAIN MAP LOOKUP IS WRONG HERE, unlike for the two datetime
tables where `datetime.tokens/prefix-safe?` holds. `datebsearch`
compares with `strncmp(key, token, TOKMAXLEN)` and TOKMAXLEN is 10,
and FIVE keys in this table are exactly ten characters:

  microsecon  millennium  millisecon  timezone_h  timezone_m

so each is a genuine PREFIX. Checked on the oracle:
`'1 millenniumXYZ'::interval` is 1000 years and
`'2 microsecondsXX'::interval` is 00:00:00.000002.

FOUR UNITS RESOLVE HERE AND ARE STILL ERRORS. `DTK_QUARTER` and the
three timezone units have no arm in `DecodeInterval`'s switch
(datetime.c:3570-3650), so they reach its `default` and raise
22007. `'1 qtr'`, `'1 quarter'`, `'1 timezone'`,
`'1 timezone_hour'` and `'1 timezone_minute'` are all invalid
intervals even though the lookup succeeds. A table-driven decoder
that trusts this table would accept every one of them.

GENERATED -- the extraction command is in the comment below.
raw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close