Class: RecipeScrapers::Scraper
- Inherits:
-
Object
- Object
- RecipeScrapers::Scraper
- Includes:
- ParsedFields
- Defined in:
- lib/recipe_scrapers/scraper.rb,
lib/recipe_scrapers/scraper/parsed_fields.rb
Overview
Reads the recipe fields of one page. parse builds one and returns its #to_recipe, so an application gets a Models::Recipe and never the page itself. A field the page does not publish is nil.
Every field reads the site declaration first, then the schema.org recipe, then OpenGraph where that applies. A site that needs code subclasses Scraper, names its host with Scraper.host and registers with Registry.register_class.
Direct Known Subclasses
RecipeScrapers::Sites::Com::SamsungFood, RecipeScrapers::Sites::De::Chefkoch, RecipeScrapers::Sites::Fr::MadameLeFigaro
Defined Under Namespace
Modules: ParsedFields
Constant Summary collapse
- CONTRACT =
Every field a scraper answers, in the order #to_h returns them.
(Models::Recipe.members - %i[url]).freeze
Instance Attribute Summary collapse
-
#document ⇒ Nokogiri::HTML5::Document
readonly
The parsed page, for a subclass that reads it directly.
-
#url ⇒ String
readonly
The address the page was read from.
Class Method Summary collapse
-
.declared_host ⇒ String?
Names the host a subclass reads, or returns it.
-
.host(value = nil) ⇒ String?
Names the host a subclass reads, or returns it.
-
.schema_reader(value = nil) ⇒ Class<Sources::SchemaOrg>
Sets the class that reads the schema.org recipe for a subclass, or returns it.
Instance Method Summary collapse
-
#author ⇒ String?
The author, several joined with ", ".
-
#canonical_url ⇒ String
The canonical link of the page resolved against Models::Recipe#url, or the url itself when the page has none.
-
#category ⇒ String?
The course, such as "Dessert".
-
#cook_time ⇒ Integer?
The cooking time in minutes.
-
#cooking_method ⇒ String?
How it is cooked, such as "Baking".
-
#cuisine ⇒ String?
The cuisine, such as "Italian".
-
#description ⇒ String?
A sentence or short paragraph about the recipe.
-
#dietary_restrictions ⇒ Array<String>?
The diets the recipe suits, such as "VeganDiet".
-
#equipment ⇒ Array<String>?
The tools a site declares, schema.org has none.
-
#host ⇒ String
The host of Models::Recipe#url, without "www.".
-
#image ⇒ String?
The address of the main image, resolved against Models::Recipe#url.
-
#ingredient_groups ⇒ Array<Models::IngredientGroup>?
The ingredients in the groups the page shows, such as "For the sauce".
-
#ingredients ⇒ Array<String>?
The ingredient lines as the page writes them.
-
#initialize(html, url:, declaration: nil) ⇒ Scraper
constructor
A new instance of Scraper.
-
#instructions ⇒ String?
The steps joined with newlines.
-
#instructions_list ⇒ Array<String>?
The steps, with section headings as their own entries.
-
#keywords ⇒ Array<String>?
The keywords, without facet entries such as "diet: vegan".
-
#language ⇒ String?
The language tag of the recipe, such as "en-US".
-
#links ⇒ Array<String>
The href of every link on the page, as written.
-
#nutrients ⇒ Hash{String => String}?
The nutrition facts as the page publishes them, keyed by the schema.org property name.
-
#prep_time ⇒ Integer?
The preparation time in minutes.
-
#ratings ⇒ Float?
The average rating.
-
#ratings_count ⇒ Integer?
How many ratings the average is made of.
-
#recipe? ⇒ Boolean
private
Whether the page has a title and ingredients.
-
#site_name ⇒ String?
The name of the website.
-
#title ⇒ String?
The name of the recipe.
-
#to_h ⇒ Hash{Symbol => Object}
private
Every field of CONTRACT with its value.
-
#to_recipe ⇒ Models::Recipe
private
Reads every field once and returns them without the page.
-
#total_time ⇒ Integer?
The total time in minutes.
-
#yields ⇒ String?
How much it makes, as "4 servings" or "12 items".
Methods included from ParsedFields
included, #parsed_ingredients, #parsed_nutrients
Constructor Details
#initialize(html, url:, declaration: nil) ⇒ Scraper
Returns a new instance of Scraper.
64 65 66 67 68 69 70 71 |
# File 'lib/recipe_scrapers/scraper.rb', line 64 def initialize(html, url:, declaration: nil) @url = url @declaration = declaration @document = Nokogiri::HTML5(Text.recode(html, declaration&.encoding_name)) @schema = self.class.schema_reader.new(Sources::JsonLd.new(@document), Sources::Microdata.new(@document)) @open_graph = Sources::OpenGraph.new(@document) @declared = Sources::Declared.new(@document, declaration) end |
Instance Attribute Details
#document ⇒ Nokogiri::HTML5::Document (readonly)
Returns the parsed page, for a subclass that reads it directly.
59 60 61 |
# File 'lib/recipe_scrapers/scraper.rb', line 59 def document @document end |
#url ⇒ String (readonly)
Returns the address the page was read from.
56 57 58 |
# File 'lib/recipe_scrapers/scraper.rb', line 56 def url @url end |
Class Method Details
.declared_host ⇒ String?
Names the host a subclass reads, or returns it.
43 44 45 46 |
# File 'lib/recipe_scrapers/scraper.rb', line 43 def host(value = nil) @declared_host = value if value @declared_host end |
.host(value = nil) ⇒ String?
Names the host a subclass reads, or returns it.
38 39 40 41 |
# File 'lib/recipe_scrapers/scraper.rb', line 38 def host(value = nil) @declared_host = value if value @declared_host end |
.schema_reader(value = nil) ⇒ Class<Sources::SchemaOrg>
Sets the class that reads the schema.org recipe for a subclass, or returns it.
49 50 51 52 |
# File 'lib/recipe_scrapers/scraper.rb', line 49 def schema_reader(value = nil) @schema_reader = value if value @schema_reader || Sources::SchemaOrg end |
Instance Method Details
#author ⇒ String?
Returns the author, several joined with ", ".
91 92 |
# File 'lib/recipe_scrapers/scraper.rb', line 91 def = @declared.text(:author) || @schema. # (see Models::Recipe#site_name) |
#canonical_url ⇒ String
Returns the canonical link of the page resolved against Models::Recipe#url, or the url itself when the page has none.
79 80 81 82 83 84 85 86 |
# File 'lib/recipe_scrapers/scraper.rb', line 79 def canonical_url link = document.at_css('link[rel="canonical"][href]') return url unless link URI.join(url, link["href"]).to_s rescue URI::Error url end |
#category ⇒ String?
Returns the course, such as "Dessert".
99 100 |
# File 'lib/recipe_scrapers/scraper.rb', line 99 def category = @declared.text(:category) || @schema.category # (see Models::Recipe#cuisine) |
#cook_time ⇒ Integer?
Returns the cooking time in minutes.
122 123 |
# File 'lib/recipe_scrapers/scraper.rb', line 122 def cook_time = Parsers::Durations.minutes(@declared.text(:cook_time)) || @schema.cook_time # (see Models::Recipe#prep_time) |
#cooking_method ⇒ String?
Returns how it is cooked, such as "Baking".
103 104 |
# File 'lib/recipe_scrapers/scraper.rb', line 103 def cooking_method = @declared.text(:cooking_method) || @schema.cooking_method # (see Models::Recipe#keywords) |
#cuisine ⇒ String?
Returns the cuisine, such as "Italian".
101 102 |
# File 'lib/recipe_scrapers/scraper.rb', line 101 def cuisine = @declared.text(:cuisine) || @schema.cuisine # (see Models::Recipe#cooking_method) |
#description ⇒ String?
Returns a sentence or short paragraph about the recipe.
97 98 |
# File 'lib/recipe_scrapers/scraper.rb', line 97 def description = @declared.text(:description) || @schema.description || @open_graph.description # (see Models::Recipe#category) |
#dietary_restrictions ⇒ Array<String>?
Returns the diets the recipe suits, such as "VeganDiet".
111 112 |
# File 'lib/recipe_scrapers/scraper.rb', line 111 def dietary_restrictions = @schema.dietary_restrictions # (see Models::Recipe#ratings) |
#equipment ⇒ Array<String>?
Returns the tools a site declares, schema.org has none.
107 108 |
# File 'lib/recipe_scrapers/scraper.rb', line 107 def equipment = @declared.rows(:equipment) # (see Models::Recipe#nutrients) |
#host ⇒ String
Returns the host of Models::Recipe#url, without "www.".
74 75 76 |
# File 'lib/recipe_scrapers/scraper.rb', line 74 def host URI.parse(url).host.to_s.delete_prefix("www.") end |
#image ⇒ String?
Returns the address of the main image, resolved against Models::Recipe#url.
148 149 150 151 152 153 |
# File 'lib/recipe_scrapers/scraper.rb', line 148 def image address = @declared.text(:image) || @schema.image || @open_graph.image address && URI.join(url, address[EMBEDDED_ADDRESS, :address] || address).to_s rescue URI::Error address end |
#ingredient_groups ⇒ Array<Models::IngredientGroup>?
The ingredients in the groups the page shows, such as "For the sauce". The groups come from the HTML when the page marks them up, otherwise from the heading lines of the ingredient list. A page without groups gives one group with a nil purpose.
138 139 140 141 142 143 144 145 |
# File 'lib/recipe_scrapers/scraper.rb', line 138 def ingredient_groups lines = ingredients return nil if lines.nil? marked = marked_groups(lines) groups = marked.any?(&:purpose) ? marked : Sources::IngredientGroups.from_sections(ingredient_sections) with_parsed_ingredients(groups) end |
#ingredients ⇒ Array<String>?
The ingredient lines as the page writes them. A line with no letter and no digit, such as "*", is a separator and left out. A line that ends in a colon and has no digit, such as "For the sauce:", is a group heading and goes to Models::Recipe#ingredient_groups instead. See Models::Recipe#parsed_ingredients.
127 128 129 130 |
# File 'lib/recipe_scrapers/scraper.rb', line 127 def ingredients lines = ingredient_sections.flat_map(&:last) lines.empty? ? nil : lines end |
#instructions ⇒ String?
Returns the steps joined with newlines.
135 |
# File 'lib/recipe_scrapers/scraper.rb', line 135 def instructions = instructions_list&.join("\n") |
#instructions_list ⇒ Array<String>?
Returns the steps, with section headings as their own entries.
133 134 |
# File 'lib/recipe_scrapers/scraper.rb', line 133 def instructions_list = @declared.rows(:instructions) || @schema.instructions_list # (see Models::Recipe#instructions) |
#keywords ⇒ Array<String>?
Returns the keywords, without facet entries such as "diet: vegan".
105 106 |
# File 'lib/recipe_scrapers/scraper.rb', line 105 def keywords = @declared.rows(:keywords) || @schema.keywords # (see Models::Recipe#equipment) |
#language ⇒ String?
Returns the language tag of the recipe, such as "en-US".
95 96 |
# File 'lib/recipe_scrapers/scraper.rb', line 95 def language = @declared.text(:language) || @schema.language || document_language # (see Models::Recipe#description) |
#links ⇒ Array<String>
Returns the href of every link on the page, as written.
156 157 158 |
# File 'lib/recipe_scrapers/scraper.rb', line 156 def links document.css("a[href]").map { |anchor| anchor["href"] } end |
#nutrients ⇒ Hash{String => String}?
The nutrition facts as the page publishes them, keyed by the schema.org property name. Nothing is parsed here. See Models::Recipe#parsed_nutrients.
109 110 |
# File 'lib/recipe_scrapers/scraper.rb', line 109 def nutrients = @schema.nutrients # (see Models::Recipe#dietary_restrictions) |
#prep_time ⇒ Integer?
Returns the preparation time in minutes.
124 |
# File 'lib/recipe_scrapers/scraper.rb', line 124 def prep_time = Parsers::Durations.minutes(@declared.text(:prep_time)) || @schema.prep_time |
#ratings ⇒ Float?
Returns the average rating.
113 114 |
# File 'lib/recipe_scrapers/scraper.rb', line 113 def = @schema. # (see Models::Recipe#ratings_count) |
#ratings_count ⇒ Integer?
Returns how many ratings the average is made of.
115 |
# File 'lib/recipe_scrapers/scraper.rb', line 115 def = @schema. |
#recipe? ⇒ Boolean
This method is part of a private API. You should avoid using this method if possible, as it may be removed or be changed in the future.
Returns whether the page has a title and ingredients.
179 180 181 |
# File 'lib/recipe_scrapers/scraper.rb', line 179 def recipe? !title.nil? && !ingredients.nil? end |
#site_name ⇒ String?
Returns the name of the website.
93 94 |
# File 'lib/recipe_scrapers/scraper.rb', line 93 def site_name = @declared.text(:site_name) || @schema.site_name || @open_graph.site_name # (see Models::Recipe#language) |
#title ⇒ String?
Returns the name of the recipe.
89 90 |
# File 'lib/recipe_scrapers/scraper.rb', line 89 def title = @declared.text(:title) || @schema.title || @open_graph.title # (see Models::Recipe#author) |
#to_h ⇒ Hash{Symbol => Object}
This method is part of a private API. You should avoid using this method if possible, as it may be removed or be changed in the future.
Returns every field of CONTRACT with its value.
163 164 165 |
# File 'lib/recipe_scrapers/scraper.rb', line 163 def to_h CONTRACT.to_h { |field| [field, public_send(field)] } end |
#to_recipe ⇒ Models::Recipe
This method is part of a private API. You should avoid using this method if possible, as it may be removed or be changed in the future.
Reads every field once and returns them without the page.
172 173 174 |
# File 'lib/recipe_scrapers/scraper.rb', line 172 def to_recipe Models::Recipe.new(url: url, **to_h) end |