Module: Ksef::FA3::FieldChecks
- Included in:
- AdvanceChecks, AttachmentChecks, CorrectionChecks, ModelValidator, SubjectChecks, SummaryChecks
- Defined in:
- lib/ksef/fa3/field_checks.rb
Overview
Checks on a single field value: is it readable, is it the right shape, is it one of the values the schema allows.
Split out of ModelValidator because none of it knows what an invoice is — it knows what
xsd:token means and what the generated enum tables contain. ModelValidator decides
which fields to check and what to call them; this decides what "too long" and "not one
of the permitted values" actually mean.
The two subtleties worth reading before changing anything here
xsd:token collapses whitespace before its facets apply. TZnakowy and TZnakowy512
both derive from it, so the schema measures — and compares — the collapsed value.
Measuring the raw string made tier 1 stricter than the schema, which on the send path
means refusing documents KSeF admits: upstream's own ksef-pdf-generator/invoice.xml has
thirteen line names of 577 raw characters that collapse to 500 (2026-08-24).
Text can be tagged UTF-8 and not be UTF-8. CLAUDE.md warns the ambient locale is not
UTF-8, and docs/REFERENCE.md §15.1 argues mis-encoded text is the likely real-world
rejection. String#strip raises Encoding::CompatibilityError on such a value, out of
this gem's hierarchy entirely — so encoding is checked before anything reads the string.
Constant Summary collapse
- SHORT_TEXT =
TZnakowyandTZnakowy512: bothminLength="1", differing only in the ceiling. A value that is present but empty is a schema violation, not an absent value. They live here rather than on ModelValidator because every module that measures a field includes this one. 256- LONG_TEXT =
512- BUYER_ID_TEXT =
IDNabywcyrestrictsTZnakowy50further, to 32. 32- COLLAPSE =
xsd:token'swhiteSpace="collapse". /\s+/- FORBIDDEN_IN_XML =
Outside XML 1.0's
Charproduction: the C0 controls except tab, newline and carriage return, plus the two plane-0 noncharacters. Not the discouraged characters of §15.1 — DocumentValidator handles those — but characters no XML document may contain at all.What happens without this check varies by character, which is worth knowing before anyone tests the guard with the wrong one. Measured through
#to_xml: U+0000 alone raises a bareArgumentError: string contains null bytefrom inside the serializer. The rest — U+0001, U+0008, U+001F, U+FFFE, U+FFFF — are written out raw, producing a document only tier 2 rejects (PCDATA invalid Char value 1). U+000B and U+000C never arrive: Ruby's\scovers them, so Ksef::FA3::Formatting.text's collapse has already eaten them. The check earns its place either way — it is what stops non-well-formed output — but only NUL escapes this gem's error hierarchy. /[\x00-\x08\x0B\x0C\x0E-\x1F]/
Class Method Summary collapse
-
.utf8?(string) ⇒ Boolean
Two different failures, and only one of them used to be caught. A string can be invalid — bytes that decode as nothing — or validly encoded in something that is not UTF-8.
Class Method Details
.utf8?(string) ⇒ Boolean
Two different failures, and only one of them used to be caught. A string can be
invalid — bytes that decode as nothing — or validly encoded in something that is
not UTF-8. String#valid_encoding? answers true for the second, so a
Windows-1250 or ISO-8859-2 name — exactly what a Polish ERP emits — passed every
guard in this gem and then raised Encoding::CompatibilityError out of #errors,
#to_xml and Ksef::Client#send_invoice, from the tier whose contract is to report.
ASCII-only text in another encoding is accepted: its bytes are UTF-8, and refusing
a File.binread of an all-ASCII document would be pedantry rather than safety.
ASCII-8BIT is treated as unlabelled rather than as another encoding. It is what
File.binread and most socket reads produce, and it asserts nothing about the bytes
— so if they happen to form valid UTF-8, reading them as UTF-8 is unambiguous rather
than a guess. A named non-UTF-8 encoding is a statement, and overriding it would
be exactly the guessing this gem refuses to do: Windows-1250 "Łódź" begins A3,
a UTF-8 continuation byte, so it fails this and is reported instead.
71 72 73 74 75 76 |
# File 'lib/ksef/fa3/field_checks.rb', line 71 def self.utf8?(string) return string.valid_encoding? if string.encoding == Encoding::UTF_8 return string.dup.force_encoding(Encoding::UTF_8).valid_encoding? if string.encoding == Encoding::BINARY string.ascii_only? end |