The leaf decoders of PostgreSQL's datetime parser: the small functions
DecodeDateTime and DecodeTimeOnly call to turn ONE lexed field
into numbers. They are separated from the state machine because each
is independently testable and several have rules no one would guess.
SIGN CONVENTION. PostgreSQL carries a zone offset internally as
SECONDS WEST of Greenwich -- DecodeTimezone computes seconds east
and then stores *tzp = -tz (datetime.c:3060). The abbreviation
table, and tznames/Default that it comes from, are SECONDS EAST.
Everything here is named for which one it is, because a sign error
between them is a silent double offset rather than a failure.
ERRORS are ex-info carrying ::dterr, mirroring the C's DTERR_*
returns. The entry points map those to SQLSTATEs; nothing here knows
about SQL.
The leaf decoders of PostgreSQL's datetime parser: the small functions `DecodeDateTime` and `DecodeTimeOnly` call to turn ONE lexed field into numbers. They are separated from the state machine because each is independently testable and several have rules no one would guess. SIGN CONVENTION. PostgreSQL carries a zone offset internally as SECONDS WEST of Greenwich -- `DecodeTimezone` computes seconds east and then stores `*tzp = -tz` (datetime.c:3060). The abbreviation table, and `tznames/Default` that it comes from, are SECONDS EAST. Everything here is named for which one it is, because a sign error between them is a silent double offset rather than a failure. ERRORS are `ex-info` carrying `::dterr`, mirroring the C's DTERR_* returns. The entry points map those to SQLSTATEs; nothing here knows about SQL.
(date2j year month day)date2j (datetime.c:286-308): Gregorian date to Julian day number.
Transcribed with its integer arithmetic intact. The 7834 * month / 256 is a trick for the 30.6-day month step and only works with
truncating integer division -- quot, not /. The function is valid
for any year including negative ones, which is how BC dates work.
`date2j` (datetime.c:286-308): Gregorian date to Julian day number. Transcribed with its integer arithmetic intact. The `7834 * month / 256` is a trick for the 30.6-day month step and only works with truncating integer division -- `quot`, not `/`. The function is valid for any year including negative ones, which is how BC dates work.
(decode-date s fmask {:keys [date-order] :as opts})DecodeDate (datetime.c:2694-2776): a whole :date FIELD -- the
thing the lexer tagged, like 2001-02-03 or 10-feb-1997.
It re-splits the field into alnum runs and then makes TWO passes,
and the two-pass structure is the point: every text run is resolved
first, so that when the numbers are placed the decoder already knows
whether a month was named. Without that, '10-Feb-1997' and
'Feb-10-1997' could not both work -- the first number is a day in
one and a month in the other, and only the text pass can tell them
apart before the numbers are read.
A text run may ONLY be a month or an ignorable token. A weekday inside a date field is a format error, where a weekday as a separate field is silently dropped.
Returns {:tm :fields :two-digits? :text-month?}, or throws.
`DecodeDate` (datetime.c:2694-2776): a whole `:date` FIELD -- the
thing the lexer tagged, like `2001-02-03` or `10-feb-1997`.
It re-splits the field into alnum runs and then makes TWO passes,
and the two-pass structure is the point: every text run is resolved
first, so that when the numbers are placed the decoder already knows
whether a month was named. Without that, `'10-Feb-1997'` and
`'Feb-10-1997'` could not both work -- the first number is a day in
one and a month in the other, and only the text pass can tell them
apart before the numbers are read.
A text run may ONLY be a month or an ignorable token. A weekday
inside a date field is a format error, where a weekday as a separate
field is silently dropped.
Returns `{:tm :fields :two-digits? :text-month?}`, or throws.(decode-number s tm fmask {:keys [text-month? date-order two-digits?]})DecodeNumber (datetime.c:2778-2910): ONE plain number field, placed
according to what is already known and to DateStyle.
Takes the tm so far and returns an updated one:
{:tm :tmask :two-digits? :kind}. :kind is present only when the
field was handed on to decode-number-field, which is the one path
that can set a TIME from what looked like a date field.
date-order is DateStyle's field order: :ymd, :dmy or :mdy.
The day-of-year rule comes FIRST, before the switch: three
characters, a year already set and nothing else, value 1..366. That
is why '1997 038' is 7 February, and why the rule cannot be folded
into the :year-only case below it. Note it claims MONTH and DAY in
its tmask as well as DOY -- it has not set them, but ValidateDate
will, and claiming them is what stops a later field also setting
them.
text-month? changes the answer in two places. With a text month
already seen, a 3+ digit number is the YEAR rather than the day --
'Feb 10 1997'. And in the YEAR|MONTH case it can RETROACTIVELY
reinterpret: if the year already set came from two digits and this
field has three or more, the earlier value was really the day, so
the two swap and two-digits? is cleared. '08-Jan-99' reaches
1999-01-08 that way.
It does NOT rescue '99-Jan-08', and I claimed it did before
checking. The swap needs the LATER field to be the long one; with
the long field first, 99 is placed as the day and the oracle agrees
that the result is out of range:
select '99-Jan-08'::date => date/time field value out of range
`DecodeNumber` (datetime.c:2778-2910): ONE plain number field, placed
according to what is already known and to DateStyle.
Takes the `tm` so far and returns an updated one:
`{:tm :tmask :two-digits? :kind}`. `:kind` is present only when the
field was handed on to `decode-number-field`, which is the one path
that can set a TIME from what looked like a date field.
`date-order` is DateStyle's field order: `:ymd`, `:dmy` or `:mdy`.
The day-of-year rule comes FIRST, before the switch: three
characters, a year already set and nothing else, value 1..366. That
is why `'1997 038'` is 7 February, and why the rule cannot be folded
into the `:year`-only case below it. Note it claims MONTH and DAY in
its tmask as well as DOY -- it has not set them, but `ValidateDate`
will, and claiming them is what stops a later field also setting
them.
`text-month?` changes the answer in two places. With a text month
already seen, a 3+ digit number is the YEAR rather than the day --
`'Feb 10 1997'`. And in the YEAR|MONTH case it can RETROACTIVELY
reinterpret: if the year already set came from two digits and this
field has three or more, the earlier value was really the day, so
the two swap and `two-digits?` is cleared. `'08-Jan-99'` reaches
1999-01-08 that way.
It does NOT rescue `'99-Jan-08'`, and I claimed it did before
checking. The swap needs the LATER field to be the long one; with
the long field first, 99 is placed as the day and the oracle agrees
that the result is out of range:
select '99-Jan-08'::date => date/time field value out of range(decode-number-field s fmask)DecodeNumberField (datetime.c:2912-3005): a run-together number
like 19970210, 040506 or 19970210.5.
Returns {:kind :date|:time :tm {…} :tmask #{…} :two-digits? bool},
or throws. fmask decides WHICH it can be, so the same digits mean
different things in different positions -- that is the function's
whole job, and it is why '040506' is a date to DecodeDateTime and
a time to DecodeTimeOnly.
A decimal point changes everything. The fraction is taken, the string
is TRUNCATED at the point (*(cp) = 0, datetime.c:2938) and len
recomputed -- and because the date branch is an else if on that
same test, a number carrying a point can only ever be a TIME. So
040506.5 is a time, while bare 040506 depends on context.
The date split is positional: the last two characters are the day,
the two before that the month, everything left the year. (len - 4) == 2 means the year part was two characters, which sets
two-digits? and brings the 1970-2069 rule into play.
`DecodeNumberField` (datetime.c:2912-3005): a run-together number
like `19970210`, `040506` or `19970210.5`.
Returns `{:kind :date|:time :tm {…} :tmask #{…} :two-digits? bool}`,
or throws. `fmask` decides WHICH it can be, so the same digits mean
different things in different positions -- that is the function's
whole job, and it is why `'040506'` is a date to `DecodeDateTime` and
a time to `DecodeTimeOnly`.
A decimal point changes everything. The fraction is taken, the string
is TRUNCATED at the point (`*(cp) = 0`, datetime.c:2938) and `len`
recomputed -- and because the date branch is an `else if` on that
same test, a number carrying a point can only ever be a TIME. So
`040506.5` is a time, while bare `040506` depends on context.
The date split is positional: the last two characters are the day,
the two before that the month, everything left the year. `(len - 4)
== 2` means the year part was two characters, which sets
`two-digits?` and brings the 1970-2069 rule into play.(decode-time s)DecodeTimeCommon/DecodeTime (datetime.c:2583-2683). A :time
field to {:hour :min :sec :usec}.
hh:mm.ff is NOT hours, minutes and a fraction of a minute: the C
shifts the fields down (tm_sec = tm_min; tm_min = tm_hour; tm_hour = 0, datetime.c:2622-2628), so 04:05.5 is four MINUTES and
five and a half seconds. That is for DecodeInterval's benefit and
it applies here all the same.
Seconds may be exactly 60 (> SECS_PER_MINUTE is the failure, not
>=): 23:59:60 is a legal input, normalized afterwards. Minutes
may not -- MINS_PER_HOUR - 1 is the limit there.
The hour has NO upper bound in this function. 25:00:00 and
24:00:00 both decode fine and are rejected, or not, by the caller:
DecodeTimeOnly allows exactly 24:00:00 and tm2time wraps nothing.
Putting the hour limit here would reject 24:00:00, which is legal.
`DecodeTimeCommon`/`DecodeTime` (datetime.c:2583-2683). A `:time`
field to `{:hour :min :sec :usec}`.
`hh:mm.ff` is NOT hours, minutes and a fraction of a minute: the C
shifts the fields down (`tm_sec = tm_min; tm_min = tm_hour;
tm_hour = 0`, datetime.c:2622-2628), so `04:05.5` is four MINUTES and
five and a half seconds. That is for `DecodeInterval`'s benefit and
it applies here all the same.
Seconds may be exactly 60 (`> SECS_PER_MINUTE` is the failure, not
`>=`): `23:59:60` is a legal input, normalized afterwards. Minutes
may not -- `MINS_PER_HOUR - 1` is the limit there.
The hour has NO upper bound in this function. `25:00:00` and
`24:00:00` both decode fine and are rejected, or not, by the caller:
`DecodeTimeOnly` allows exactly 24:00:00 and `tm2time` wraps nothing.
Putting the hour limit here would reject `24:00:00`, which is legal.(decode-timezone s)DecodeTimezone. A numeric zone field (+05, -08:30, +0530,
+05:30:30) to SECONDS WEST of Greenwich.
The bare-digits rule is the part that surprises: with no colon, a
string LONGER THAN 3 characters is split as hhmm, and a shorter one
is hours. So +05 is five hours, +0530 is five and a half -- and
+100 is ONE hour, because it is four characters and splits to
hr=1, min=0. Length, not value.
The trailing check (*cp != '\0') happens AFTER the arithmetic but
still rejects: +05.5 is a format error, not five and a half hours.
The lexer's greedy run over . : - exists precisely so that the
whole of +05.5 arrives here to be refused, rather than .5
escaping into the fraction.
`DecodeTimezone`. A numeric zone field (`+05`, `-08:30`, `+0530`, `+05:30:30`) to SECONDS WEST of Greenwich. The bare-digits rule is the part that surprises: with no colon, a string LONGER THAN 3 characters is split as hhmm, and a shorter one is hours. So `+05` is five hours, `+0530` is five and a half -- and `+100` is ONE hour, because it is four characters and splits to hr=1, min=0. Length, not value. The trailing check (`*cp != '\0'`) happens AFTER the arithmetic but still rejects: `+05.5` is a format error, not five and a half hours. The lexer's greedy run over `. : -` exists precisely so that the whole of `+05.5` arrives here to be refused, rather than `.5` escaping into the fraction.
(dterr kind)A DTERR_* return, as a throw. The C returns these as ints and every caller checks; a throw is the same control flow with less to forget.
A DTERR_* return, as a throw. The C returns these as ints and every caller checks; a throw is the same control flow with less to forget.
(j2date jd)j2date (datetime.c:311-333): Julian day number back to
[year month day].
The C declares julian, quad and extra as UNSIGNED int and
relies on 32-bit wraparound nowhere in the valid range -- but the
intermediate (julian - quad * 146097) * 4 + 3 would overflow a
signed 32-bit int for large inputs, which is why they are unsigned.
Clojure longs are 64-bit, so the arithmetic is exact and no masking
is needed; the division and % stay truncating to match.
`j2date` (datetime.c:311-333): Julian day number back to `[year month day]`. The C declares `julian`, `quad` and `extra` as UNSIGNED int and relies on 32-bit wraparound nowhere in the valid range -- but the intermediate `(julian - quad * 146097) * 4 + 3` would overflow a signed 32-bit int for large inputs, which is why they are unsigned. Clojure longs are 64-bit, so the arithmetic is exact and no masking is needed; the division and `%` stay truncating to match.
(j2day jd)j2day (datetime.c:338-352): Julian day to day-of-week, 0..6 =
Sun..Sat.
`j2day` (datetime.c:338-352): Julian day to day-of-week, 0..6 = Sun..Sat.
(leap-year? y)isleap (datetime.h:271). Applied to the PROLEPTIC year as
PostgreSQL stores it, where 1 BC is year 0 -- so year 0 is a leap
year and 0000-02-29 BC is a real date.
`isleap` (datetime.h:271). Applied to the PROLEPTIC year as PostgreSQL stores it, where 1 BC is year 0 -- so year 0 is a leap year and `0000-02-29 BC` is a real date.
MAX_TZDISP_HOUR (timestamp.h). The largest zone displacement PostgreSQL accepts -- not 24, and not 14 either.
MAX_TZDISP_HOUR (timestamp.h). The largest zone displacement PostgreSQL accepts -- not 24, and not 14 either.
(validate-date {:keys [year mon mday yday] :as tm}
fields
{:keys [julian? two-digits? bc?]})ValidateDate. Applies the year rules, resolves a day-of-year, and
range-checks -- in that order, because each step needs the previous
one's answer.
fields is the set of date fields that were actually SET, which is
the fmask test in the C; a check that a field is in range must not
run on a field nobody supplied.
Returns the corrected tm, or throws.
Two rules here are load-bearing and neither is obvious:
There is no year zero in AD/BC notation, so a non-Julian year <= 0 is
an overflow -- but isjulian SKIPS the check entirely, because a
Julian day legitimately produces year 0 and below.
The 2-digit year rule (<70 => +2000, <100 => +1900) is applied only
when is2digits AND NOT bc: the branches are an if/else chain with
bc first. So '70-01-01' is 1970 and '70-01-01 BC' is 70 BC, not
1970 BC.
`ValidateDate`. Applies the year rules, resolves a day-of-year, and range-checks -- in that order, because each step needs the previous one's answer. `fields` is the set of date fields that were actually SET, which is the `fmask` test in the C; a check that a field is in range must not run on a field nobody supplied. Returns the corrected `tm`, or throws. Two rules here are load-bearing and neither is obvious: There is no year zero in AD/BC notation, so a non-Julian year <= 0 is an overflow -- but `isjulian` SKIPS the check entirely, because a Julian day legitimately produces year 0 and below. The 2-digit year rule (<70 => +2000, <100 => +1900) is applied only when `is2digits` AND NOT `bc`: the branches are an if/else chain with `bc` first. So `'70-01-01'` is 1970 and `'70-01-01 BC'` is 70 BC, not 1970 BC.
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |