ParseDateTime (datetime.c:753-977): the datetime LEXER.
It splits a literal into fields and tags each with a TYPE, without
deciding what any of them means. That split -- lex, then interpret --
is the whole reason this exists rather than a regex: the meaning of a
field depends on the other fields present and on DateStyle, and a
pattern cannot carry that. A bare three-digit number is a day-of-year
if a year is already set and a month otherwise; 01/02/03 is three
different dates; Mon Feb 10 17:32:01 1997 PST has a word that is a
month, a word that is a weekday to discard, and a word that is a zone.
Six field types, and the doc comment at datetime.c:740-752 is worth repeating because it is the only place that says the quiet part -- several of them hold things their names do not suggest:
:number digits and possibly a decimal point. ALSO holds a date:
yy.ddd.
:string text with no digits or punctuation. ALSO holds months
(january) and zone abbreviations (pst).
:date digits with two delimiters, or digits and text. ALSO
holds zone NAMES: america/new_york, gmt-8.
:time digits with colon delimiters, possibly a decimal point.
:tz a leading + or - then digits (and it eats :, .
and - as well -- see the note on that below).
:special a leading + or - then text.
Alpha runs are lowercased; digits and punctuation are not, and the
fold goes through tokens/ascii-lower rather than
clojure.string/lower-case.
That assumes a UTF-8 database, and the assumption is worth stating.
pg_tolower also folds bytes with the high bit set when isupper
says to, and the C's isalpha likewise respects LC_CTYPE -- but in
a UTF-8 locale every single byte over 0x7F is false for both, so the
ASCII-only fold agrees exactly. In a single-byte database (LATIN1
with a matching LC_CTYPE) PostgreSQL would lex 1-Été-2000 and we
reject it. pg-datahike serves UTF-8 only, so that case is
unreachable here; it is recorded because the reason it is safe is
not visible in the code.
Punctuation that is not part of a field is DISCARDED as a delimiter
(datetime.c:929-934). So a double-quoted zone in a literal is not a
quoted anything: '2000-01-01 12:00 "PST"' lexes exactly like
… PST, because " is punctuation. There is no quoted-zone concept
at the literal level at all.
Returns a vector of {:text :type}, or throws bad-format.
`ParseDateTime` (datetime.c:753-977): the datetime LEXER.
It splits a literal into fields and tags each with a TYPE, without
deciding what any of them means. That split -- lex, then interpret --
is the whole reason this exists rather than a regex: the meaning of a
field depends on the other fields present and on DateStyle, and a
pattern cannot carry that. A bare three-digit number is a day-of-year
if a year is already set and a month otherwise; `01/02/03` is three
different dates; `Mon Feb 10 17:32:01 1997 PST` has a word that is a
month, a word that is a weekday to discard, and a word that is a zone.
Six field types, and the doc comment at datetime.c:740-752 is worth
repeating because it is the only place that says the quiet part --
several of them hold things their names do not suggest:
:number digits and possibly a decimal point. ALSO holds a date:
`yy.ddd`.
:string text with no digits or punctuation. ALSO holds months
(`january`) and zone abbreviations (`pst`).
:date digits with two delimiters, or digits and text. ALSO
holds zone NAMES: `america/new_york`, `gmt-8`.
:time digits with colon delimiters, possibly a decimal point.
:tz a leading `+` or `-` then digits (and it eats `:`, `.`
and `-` as well -- see the note on that below).
:special a leading `+` or `-` then text.
Alpha runs are lowercased; digits and punctuation are not, and the
fold goes through `tokens/ascii-lower` rather than
`clojure.string/lower-case`.
That assumes a UTF-8 database, and the assumption is worth stating.
`pg_tolower` also folds bytes with the high bit set when `isupper`
says to, and the C's `isalpha` likewise respects LC_CTYPE -- but in
a UTF-8 locale every single byte over 0x7F is false for both, so the
ASCII-only fold agrees exactly. In a single-byte database (LATIN1
with a matching LC_CTYPE) PostgreSQL would lex `1-Été-2000` and we
reject it. pg-datahike serves UTF-8 only, so that case is
unreachable here; it is recorded because the reason it is safe is
not visible in the code.
Punctuation that is not part of a field is DISCARDED as a delimiter
(datetime.c:929-934). So a double-quoted zone in a literal is not a
quoted anything: `'2000-01-01 12:00 "PST"'` lexes exactly like
`… PST`, because `"` is punctuation. There is no quoted-zone concept
at the literal level at all.
Returns a vector of `{:text :type}`, or throws `bad-format`.(bad-format)DTERR_BAD_FORMAT. The caller turns it into 22007 with the type name and the original string, which only it knows.
DTERR_BAD_FORMAT. The caller turns it into 22007 with the type name and the original string, which only it knows.
ParseDateTime's work-buffer size, which is NOT a constant -- it is
chosen by each CALLER, so the same literal can be too long for one
type and fit in another:
timestamp, timestamptz MAXDATELEN + MAXDATEFIELDS = 153 (timestamp.c:182) date, time, timetz MAXDATELEN + 1 = 129 (date.c:127) interval 256, written as a literal (timestamp.c:930)
The buffer is a real input-length limit and it is observable:
'2001-02-03 04:05:06.' || repeat('0',132) is a valid timestamp and
one more zero is 22007. Without it the port silently accepted
arbitrarily long input and returned an answer.
`ParseDateTime`'s work-buffer size, which is NOT a constant -- it is
chosen by each CALLER, so the same literal can be too long for one
type and fit in another:
timestamp, timestamptz MAXDATELEN + MAXDATEFIELDS = 153
(timestamp.c:182)
date, time, timetz MAXDATELEN + 1 = 129 (date.c:127)
interval 256, written as a literal (timestamp.c:930)
The buffer is a real input-length limit and it is observable:
`'2001-02-03 04:05:06.' || repeat('0',132)` is a valid timestamp and
one more zero is 22007. Without it the port silently accepted
arbitrarily long input and returned an answer.MAXDATEFIELDS (datetime.h:202). More fields than this is a format error, not a truncation.
MAXDATEFIELDS (datetime.h:202). More fields than this is a format error, not a truncation.
(tokenize input)(tokenize input buflen)ParseDateTime. The literal as [{:text :type} …].
The scanner's branch order is PostgreSQL's and the order matters: a leading digit is examined by what FOLLOWS it, and three of the five sub-cases hinge on whether the character after a delimiter is itself a digit.
`ParseDateTime`. The literal as `[{:text :type} …]`.
The scanner's branch order is PostgreSQL's and the order matters: a
leading digit is examined by what FOLLOWS it, and three of the five
sub-cases hinge on whether the character after a delimiter is itself
a digit.cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |