Module: Canon::Xml::WhitespacePolicy
- Defined in:
- lib/canon/xml/whitespace_policy.rb
Overview
Parse-time whitespace policy: whether a character-data node survives conversion. One home for the keep/strip rules so the DOM/SAX/HTML differences are visible here instead of implied by copy-paste across conversion sites.
Document-level character data (outside the root element) can only be whitespace per the XML grammar, and the XPath data model — which canon's tree and C14N follow — has no root text children at all. Every policy therefore drops whitespace-only document-level text; engines that report it (libxml2 does not, libleptris 1.9.38+ does) stay byte-compatible through here.
Constant Summary collapse
- STRIP_ONLY =
Zero-allocation forms of the
content.strip.empty?/content.gsub(...).empty?checks these policies used to run per text node. STRIP_ONLY is exactly String#strip's set (ASCII whitespace plus null); SAX drops space/tab/CR/LF runs. /\A[\0\t\n\v\f\r ]*\z/- SAX_DROPPED =
/\A[ \t\r\n]*\z/- HTML_WHITESPACE_SENSITIVE_TAGS =
HTML conversion rule: whitespace-only text is dropped except in whitespace-sensitive elements (pre/code/textarea/script/style), between inline siblings (semantically significant), and when it carries NBSP (U+00A0 — never insignificant; strip is ASCII-only so it is checked explicitly).
%w[pre code textarea script style].freeze
- HTML_SENSITIVE_TAG_SET =
Set form for the hot lookup; element names arrive lowercased from the HTML parser, so the downcase fallback only allocates for unusual (non-lowercased) names.
HTML_WHITESPACE_SENSITIVE_TAGS.to_set
Class Method Summary collapse
-
.keep_dom_text?(content, preserve_whitespace:, element_parent: true) ⇒ Boolean
DOM conversion rule: whitespace-only text is dropped unless preserving.
- .keep_html_text?(content, parent_name:, text_node: nil) ⇒ Boolean
-
.keep_sax_text?(content, preserve_whitespace:, element_parent: true) ⇒ Boolean
SAX rule: same shape, plus CR-bearing content is always kept ( must survive parsing for C14N) — only runs of pure ASCII whitespace (space, tab, CR, LF) are dropped when not preserving.
Class Method Details
.keep_dom_text?(content, preserve_whitespace:, element_parent: true) ⇒ Boolean
DOM conversion rule: whitespace-only text is dropped unless preserving. Non-ASCII whitespace (NBSP, U+3000) survives — String#strip only removes ASCII whitespace.
NOTE: CR-only nodes are dropped on this path; the SAX rule keeps them (character references must survive for C14N).
34 35 36 37 38 39 40 41 42 43 44 |
# File 'lib/canon/xml/whitespace_policy.rb', line 34 def keep_dom_text?(content, preserve_whitespace:, element_parent: true) # The XPath data model has no root text children — drop # whitespace-only document-level text even when preserving # (engines that report it stay byte-compatible with libxml2). return false if !element_parent && content.match?(STRIP_ONLY) return true if preserve_whitespace !content.match?(STRIP_ONLY) end |
.keep_html_text?(content, parent_name:, text_node: nil) ⇒ Boolean
70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 |
# File 'lib/canon/xml/whitespace_policy.rb', line 70 def keep_html_text?(content, parent_name:, text_node: nil) return true unless content.match?(STRIP_ONLY) return true if content.include?(" ") parent_name = parent_name.to_s return true if HTML_SENSITIVE_TAG_SET.include?(parent_name) return true if HTML_SENSITIVE_TAG_SET.include?(parent_name.downcase) # Computed last: the sibling scan is O(siblings), so it must # not run for the content-bearing text nodes that fail the # whitespace-only check above. return true if text_node && Canon::Comparison::WhitespaceSensitivity.inline_whitespace_significant?(text_node) false end |
.keep_sax_text?(content, preserve_whitespace:, element_parent: true) ⇒ Boolean
SAX rule: same shape, plus CR-bearing content is always kept ( must survive parsing for C14N) — only runs of pure ASCII whitespace (space, tab, CR, LF) are dropped when not preserving.
49 50 51 52 53 54 55 56 57 |
# File 'lib/canon/xml/whitespace_policy.rb', line 49 def keep_sax_text?(content, preserve_whitespace:, element_parent: true) return false if !element_parent && content.match?(STRIP_ONLY) return true if preserve_whitespace return true if content.include?("\r") !content.match?(SAX_DROPPED) end |