The interval CARRIER, and interval_out.
A deftype, deliberately, and not a defrecord. The consolidation
plan warns that "a record's structural =/hash would split GROUP BY
and joins", and that is true -- but it is a reason to choose the
carrier's equality, not a reason to avoid a carrier. PostgreSQL
compares '1 mon' and '30 days' as EQUAL (interval_cmp_value:
a month is 30 days, a day is 24 hours) while PRINTING them
differently, so:
structural equality splits two values PostgreSQL groups; canonical TEXT cannot be the key either, for the same reason.
clojure.core/= routes a non-collection, non-Number object to
.equals, and clojure.core/hash routes a non-IHashEq,
non-Number, non-String object to .hashCode. So implementing those
two over interval_cmp_value makes the EXISTING machinery correct,
including the parts no comparator can reach: GROUP BY is
datahike's own group-by over the non-aggregate :find elements,
and DISTINCT, clojure.set/intersection, contains? on a set and
datahike's hash join are all clojure.core calls. None of them
takes a comparator; all of them are served by .hashCode.
Comparable covers ORDER BY, the comparison operators, BETWEEN and
min/max, through fns/order-cmp's existing :else (compare a b) --
so that function needs no interval branch at all. The same route
bits.clj's PgBit takes.
THE SPAN NEEDS MORE THAN 64 BITS. month * 30 + day can reach
6.4e10 days, and multiplying that by 86,400,000,000 µs overflows
int64 -- which is why the C uses INT128. compareTo therefore works
in BigInteger. hashCode does NOT: interval_hash deliberately
narrows the INT128 to its low 64 bits (int128_to_int64) before
hashing, so two spans differing by exactly 2^64 µs hash together in
PostgreSQL, and matching that is what keeps our hashing consistent
with a real server's.
toString is NOT the SQL rendering. Output goes through
types/->pg-text, which is the single funnel for ::text, array
elements, record fields, COPY, pg_dump and to_jsonb.
The interval CARRIER, and `interval_out`. A `deftype`, deliberately, and not a `defrecord`. The consolidation plan warns that "a record's structural `=`/hash would split GROUP BY and joins", and that is true -- but it is a reason to choose the carrier's equality, not a reason to avoid a carrier. PostgreSQL compares `'1 mon'` and `'30 days'` as EQUAL (`interval_cmp_value`: a month is 30 days, a day is 24 hours) while PRINTING them differently, so: structural equality splits two values PostgreSQL groups; canonical TEXT cannot be the key either, for the same reason. `clojure.core/=` routes a non-collection, non-Number object to `.equals`, and `clojure.core/hash` routes a non-`IHashEq`, non-Number, non-String object to `.hashCode`. So implementing those two over `interval_cmp_value` makes the EXISTING machinery correct, including the parts no comparator can reach: `GROUP BY` is datahike's own `group-by` over the non-aggregate `:find` elements, and `DISTINCT`, `clojure.set/intersection`, `contains?` on a set and datahike's hash join are all `clojure.core` calls. None of them takes a comparator; all of them are served by `.hashCode`. `Comparable` covers ORDER BY, the comparison operators, BETWEEN and min/max, through `fns/order-cmp`'s existing `:else (compare a b)` -- so that function needs no interval branch at all. The same route `bits.clj`'s `PgBit` takes. THE SPAN NEEDS MORE THAN 64 BITS. `month * 30 + day` can reach 6.4e10 days, and multiplying that by 86,400,000,000 µs overflows int64 -- which is why the C uses INT128. `compareTo` therefore works in `BigInteger`. `hashCode` does NOT: `interval_hash` deliberately narrows the INT128 to its low 64 bits (`int128_to_int64`) before hashing, so two spans differing by exactly 2^64 µs hash together in PostgreSQL, and matching that is what keeps our hashing consistent with a real server's. `toString` is NOT the SQL rendering. Output goes through `types/->pg-text`, which is the single funnel for `::text`, array elements, record fields, COPY, pg_dump and `to_jsonb`.
DecodeInterval (datetime.c:3364-3749) and the Adjust* helpers
(560-670).
It reads the SAME lexed fields the datetime decoders read -- it is a
fourth consumer of datetime/lex.clj, with buflen 256 rather than
153 or 129 -- and then interprets them by a different set of rules.
THE LOOP RUNS RIGHT TO LEFT (datetime.c:3397). Units FOLLOW their
values in an interval literal, so '1 day' is read as day and
then 1, with the unit setting type for the field about to be
read. Nothing else in the datetime family works this way, and
reversing it is not a detail: parsing_unit_val, AGO's
last-field-only rule and DTK_HOUR's hand-off all depend on the
direction.
The accumulator is pg_itm_in (timestamp.h:82): {:usec :mday :mon :year}, where YEAR is kept separate from MONTH during decoding and
folded in only at the end by itmin2interval. Keeping them apart is
what lets AdjustYears multiply by 10, 100 and 1000 for decade,
century and millennium without losing the int32 range.
OVERFLOW IS CHECKED AT EVERY STEP in the C, with
pg_add_s32_overflow and friends, and each failure is
DTERR_FIELD_OVERFLOW -- which interval_in then REMAPS to
DTERR_INTERVAL_OVERFLOW, SQLSTATE 22015, not 22008
(timestamp.c:932-941). Clojure longs do not overflow at int32, so
the bounds are checked explicitly; a missing check is a silently
wrapped value rather than an error.
`DecodeInterval` (datetime.c:3364-3749) and the `Adjust*` helpers
(560-670).
It reads the SAME lexed fields the datetime decoders read -- it is a
fourth consumer of `datetime/lex.clj`, with `buflen` 256 rather than
153 or 129 -- and then interprets them by a different set of rules.
THE LOOP RUNS RIGHT TO LEFT (datetime.c:3397). Units FOLLOW their
values in an interval literal, so `'1 day'` is read as `day` and
then `1`, with the unit setting `type` for the field about to be
read. Nothing else in the datetime family works this way, and
reversing it is not a detail: `parsing_unit_val`, AGO's
last-field-only rule and `DTK_HOUR`'s hand-off all depend on the
direction.
The accumulator is `pg_itm_in` (timestamp.h:82): `{:usec :mday :mon
:year}`, where YEAR is kept separate from MONTH during decoding and
folded in only at the end by `itmin2interval`. Keeping them apart is
what lets `AdjustYears` multiply by 10, 100 and 1000 for decade,
century and millennium without losing the int32 range.
OVERFLOW IS CHECKED AT EVERY STEP in the C, with
`pg_add_s32_overflow` and friends, and each failure is
DTERR_FIELD_OVERFLOW -- which `interval_in` then REMAPS to
DTERR_INTERVAL_OVERFLOW, SQLSTATE 22015, not 22008
(timestamp.c:932-941). Clojure longs do not overflow at int32, so
the bounds are checked explicitly; a missing check is a silently
wrapped value rather than an error.deltatktbl (datetime.c:187-251), PostgreSQL's INTERVAL unit table.
A THIRD token table, beside datetime/tokens.clj's two. It is not a
superset of either and the three disagree on purpose:
m is MINUTE here and MONTH in datetktbl
mm is only in datetktbl
j is only in datetktbl
And the disagreement is REACHABLE, because DecodeInterval's
DTK_STRING arm tries DecodeUnits (this table) and then FALLS BACK
to DecodeSpecial (datetktbl) -- datetime.c:3659-3663. So '1 m'
is one minute and '1 mm' is also one minute, by two different
routes. Three tables, consulted in a fixed order.
A PLAIN MAP LOOKUP IS WRONG HERE, unlike for the two datetime
tables where datetime.tokens/prefix-safe? holds. datebsearch
compares with strncmp(key, token, TOKMAXLEN) and TOKMAXLEN is 10,
and FIVE keys in this table are exactly ten characters:
microsecon millennium millisecon timezone_h timezone_m
so each is a genuine PREFIX. Checked on the oracle:
'1 millenniumXYZ'::interval is 1000 years and
'2 microsecondsXX'::interval is 00:00:00.000002.
FOUR UNITS RESOLVE HERE AND ARE STILL ERRORS. DTK_QUARTER and the
three timezone units have no arm in DecodeInterval's switch
(datetime.c:3570-3650), so they reach its default and raise
22007. '1 qtr', '1 quarter', '1 timezone',
'1 timezone_hour' and '1 timezone_minute' are all invalid
intervals even though the lookup succeeds. A table-driven decoder
that trusts this table would accept every one of them.
GENERATED -- the extraction command is in the comment below.
`deltatktbl` (datetime.c:187-251), PostgreSQL's INTERVAL unit table. A THIRD token table, beside `datetime/tokens.clj`'s two. It is not a superset of either and the three disagree on purpose: `m` is MINUTE here and MONTH in `datetktbl` `mm` is only in `datetktbl` `j` is only in `datetktbl` And the disagreement is REACHABLE, because `DecodeInterval`'s DTK_STRING arm tries `DecodeUnits` (this table) and then FALLS BACK to `DecodeSpecial` (`datetktbl`) -- datetime.c:3659-3663. So `'1 m'` is one minute and `'1 mm'` is also one minute, by two different routes. Three tables, consulted in a fixed order. A PLAIN MAP LOOKUP IS WRONG HERE, unlike for the two datetime tables where `datetime.tokens/prefix-safe?` holds. `datebsearch` compares with `strncmp(key, token, TOKMAXLEN)` and TOKMAXLEN is 10, and FIVE keys in this table are exactly ten characters: microsecon millennium millisecon timezone_h timezone_m so each is a genuine PREFIX. Checked on the oracle: `'1 millenniumXYZ'::interval` is 1000 years and `'2 microsecondsXX'::interval` is 00:00:00.000002. FOUR UNITS RESOLVE HERE AND ARE STILL ERRORS. `DTK_QUARTER` and the three timezone units have no arm in `DecodeInterval`'s switch (datetime.c:3570-3650), so they reach its `default` and raise 22007. `'1 qtr'`, `'1 quarter'`, `'1 timezone'`, `'1 timezone_hour'` and `'1 timezone_minute'` are all invalid intervals even though the lookup succeeds. A table-driven decoder that trusts this table would accept every one of them. GENERATED -- the extraction command is in the comment below.
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |