Module: Ksef::FA3::DocumentValidator
- Defined in:
- lib/ksef/fa3/document_validator.rb
Overview
Tier 1b — the admission rules KSeF applies to the bytes (docs/REFERENCE.md §15.1, DESIGN.md §7.7 as amended 2026-08-24).
These exist as a separate tier because of a shape problem that went unnoticed until the rules were pinned. §7.7 described tier 1 as a model tier — required fields, enums, checksums, dates — and four of the six rules KSeF actually applies are properties of the serialized document: a byte-order mark, the prolog's declared encoding, processing instructions, and discouraged Unicode characters. At model time there is no document, so a model tier structurally cannot see any of them.
And tier 2 cannot see them either, which is the whole justification. Upstream ships
invoice-template-fa-3-with-disallowed-unicode-characters.xml, which carries U+0087 and
U+009B and is XSD-valid against the pinned schema once its #nip# placeholders are
substituted — which spec/support/fa3_corpus.rb does on read, the pinned bytes staying
verbatim so their digests keep verifying. A schema-only client sends that invoice and
KSeF rejects it.
Runs on #to_xml's output, and must run before those bytes are hashed and encrypted —
after that point a rejection costs a round trip and a session.
Constant Summary collapse
- BOM =
EF BB BF. Legal Unicode, legal XML, and an outright rejection here. "\xEF\xBB\xBF"- PROLOG =
Only the encoding matters, and only when a prolog is present at all — the prolog is optional, but if it declares anything other than UTF-8 the document is refused.
/\A<\?xml\s[^>]*\?>/- PROLOG_ENCODING =
/encoding\s*=\s*["']([^"']+)["']/- PROCESSING_INSTRUCTION =
A processing instruction, excluding the XML declaration — which is not a PI, though it looks like one. Anchoring on a name that is not
xmlis what separates them.The target is an XML
Name, which may contain non-ASCII letters, so it cannot be matched with\w:<?źdźbło x?>is a real processing instruction that an ASCII-only class waved through — the one input a review on 2026-08-24 found that every tier passed while §15.1 says KSeF rejects it. /<\?(?!xml[\s?])([^\s?>]+)/- NON_MARKUP =
Stripped before the search above, because
<?php ... ?>written inside a comment or a CDATA section is text, not a processing instruction, and rejecting it would refuse an admissible document. Characters are not scanned this way: a discouraged character is forbidden wherever it appears, comments included. /<!--.*?-->|<!\[CDATA\[.*?\]\]>/m
- DISCOURAGED =
§15.1's discouraged characters, exactly as the pinned document lists them. Note U+0085 sits between the first two ranges and is not forbidden, and that the plane noncharacters start at plane 1 — plane 0's U+FFFE/U+FFFF are excluded from XML's
Charproduction already, so they cannot occur in a well-formed document. [ 0x7F..0x84, 0x86..0x9F, 0xFDD0..0xFDEF, *(1..16).map { |plane| ((plane << 16) | 0xFFFE)..((plane << 16) | 0xFFFF) } ].freeze
- MAX_BYTES =
1 000 000 bytes, and upstream means the decimal million rather than 2^20 — it writes "1 MB * (1 000 000 bajtów)" (docs/REFERENCE.md §6.2, §15.5).
An invoice carrying an attachment gets MAX_BYTES_WITH_ATTACHMENT instead. This comment used to say attachments "are batch-only, so 0.1 has no reason to carry the larger figure", which was true exactly as long as the model could not carry one. Once
Zalaczniklanded, a legal 2.5 MB attachment invoice was refused by tier 1b with a message asserting the document had no attachment — the "refuses legal invoices" failure §15.6 and §14.3 exist to prevent, introduced by the change that made it reachable. It is a default, not a ceiling of the format. Upstream marks the figure with an asterisk — "Jeżeli w scenariuszach biznesowych organizacji dostępne limity są niewystarczające, prosimy o kontakt z działem wsparcia KSeF" — andlimity.mdheads the same numbers "Wartość domyślna", withGET /limits/contextreturning the live values for a context. So an organisation that has negotiated a higher limit passes its own viamax_bytes:; hard-coding this as absolute rejected invoices KSeF would have accepted. 1_000_000- MAX_BYTES_WITH_ATTACHMENT =
3 MB, per the same table. Attachments are documented as batch-only (§15.5, with an offline technical-correction exception), and this gem has no batch layer — but the size a document is allowed to be is not the transport's business, and tier 1b must not refuse a document the format permits.
3_000_000- DISCOURAGED_PATTERN =
DISCOURAGED compiled into one character class.
Regexp.new( "[#{DISCOURAGED.map { |range| "\\u{#{range.first.to_s(16)}}-\\u{#{range.last.to_s(16)}}" }.join}]" ).freeze
Class Method Summary collapse
-
.default_max_bytes(attachment:) ⇒ Integer
The default ceiling for such a document.
-
.errors_for(xml, max_bytes: MAX_BYTES) ⇒ Array<Issue>
Empty when KSeF would admit these bytes.
- .valid?(xml, max_bytes: MAX_BYTES) ⇒ Boolean
Class Method Details
.default_max_bytes(attachment:) ⇒ Integer
Returns the default ceiling for such a document.
84 |
# File 'lib/ksef/fa3/document_validator.rb', line 84 def self.default_max_bytes(attachment:) = ? MAX_BYTES_WITH_ATTACHMENT : MAX_BYTES |
.errors_for(xml, max_bytes: MAX_BYTES) ⇒ Array<Issue>
Returns empty when KSeF would admit these bytes.
97 98 99 100 101 102 103 104 105 106 |
# File 'lib/ksef/fa3/document_validator.rb', line 97 def errors_for(xml, max_bytes: MAX_BYTES) xml = xml.to_s # Everything below reads the string as text, and every one of those reads raises on # invalid bytes rather than reporting them. It is also a rule in its own right: §15.1 # requires the document to *be* UTF-8, not merely to lack a byte-order mark. return [encoding_issue] unless FieldChecks.utf8?(xml) [bom_issue(xml), prolog_issue(xml), *instruction_issues(xml), *character_issues(xml), size_issue(xml, max_bytes)].compact end |
.valid?(xml, max_bytes: MAX_BYTES) ⇒ Boolean
108 |
# File 'lib/ksef/fa3/document_validator.rb', line 108 def valid?(xml, max_bytes: MAX_BYTES) = errors_for(xml, max_bytes: max_bytes).empty? |