Class: RecipeScrapers::Scraper

Inherits:
Object
  • Object
show all
Includes:
ParsedFields
Defined in:
lib/recipe_scrapers/scraper.rb,
lib/recipe_scrapers/scraper/parsed_fields.rb

Overview

Reads the recipe fields of one page. parse builds one and returns its #to_recipe, so an application gets a Models::Recipe and never the page itself. A field the page does not publish is nil.

Every field reads the site declaration first, then the schema.org recipe, then OpenGraph where that applies. A site that needs code subclasses Scraper, names its host with Scraper.host and registers with Registry.register_class.

Examples:

A site with its own schema.org reader

class Chefkoch < RecipeScrapers::Scraper
  class SchemaReader < RecipeScrapers::Sources::SchemaOrg
    def ingredients = super&.map(&:strip)
  end

  host "chefkoch.de"
  schema_reader SchemaReader
end
RecipeScrapers::Registry.register_class(Chefkoch)

Defined Under Namespace

Modules: ParsedFields

Constant Summary collapse

CONTRACT =

Every field a scraper answers, in the order #to_h returns them.

(Models::Recipe.members - %i[url]).freeze

Instance Attribute Summary collapse

Class Method Summary collapse

Instance Method Summary collapse

Methods included from ParsedFields

included, #parsed_ingredients, #parsed_nutrients

Constructor Details

#initialize(html, url:, declaration: nil) ⇒ Scraper

Returns a new instance of Scraper.

Parameters:

  • html (String) —

    the page source

  • url (String) —

    the address of the page

  • declaration (Declaration, nil) (defaults to: nil) —

    the declaration of the site



64
65
66
67
68
69
70
71
# File 'lib/recipe_scrapers/scraper.rb', line 64

def initialize(html, url:, declaration: nil)
  @url = url
  @declaration = declaration
  @document = Nokogiri::HTML5(Text.recode(html, declaration&.encoding_name))
  @schema = self.class.schema_reader.new(Sources::JsonLd.new(@document), Sources::Microdata.new(@document))
  @open_graph = Sources::OpenGraph.new(@document)
  @declared = Sources::Declared.new(@document, declaration)
end

Instance Attribute Details

#document ⇒ Nokogiri::HTML5::Document (readonly)

Returns the parsed page, for a subclass that reads it directly.

Returns:

  • (Nokogiri::HTML5::Document) —

    the parsed page, for a subclass that reads it directly



59
60
61
# File 'lib/recipe_scrapers/scraper.rb', line 59

def document
  @document
end

#url ⇒ String (readonly)

Returns the address the page was read from.

Returns:

  • (String) —

    the address the page was read from



56
57
58
# File 'lib/recipe_scrapers/scraper.rb', line 56

def url
  @url
end

Class Method Details

.declared_host ⇒ String?

Names the host a subclass reads, or returns it.

Parameters:

  • value (String, nil)

Returns:

  • (String, nil)


43
44
45
46
# File 'lib/recipe_scrapers/scraper.rb', line 43

def host(value = nil)
  @declared_host = value if value
  @declared_host
end

.host(value = nil) ⇒ String?

Names the host a subclass reads, or returns it.

Parameters:

  • value (String, nil) (defaults to: nil)

Returns:

  • (String, nil)


38
39
40
41
# File 'lib/recipe_scrapers/scraper.rb', line 38

def host(value = nil)
  @declared_host = value if value
  @declared_host
end

.schema_reader(value = nil) ⇒ Class<Sources::SchemaOrg>

Sets the class that reads the schema.org recipe for a subclass, or returns it.

Parameters:

Returns:



49
50
51
52
# File 'lib/recipe_scrapers/scraper.rb', line 49

def schema_reader(value = nil)
  @schema_reader = value if value
  @schema_reader || Sources::SchemaOrg
end

Instance Method Details

#author ⇒ String?

Returns the author, several joined with ", ".

Returns:

  • (String, nil) —

    the author, several joined with ", "



91
92
# File 'lib/recipe_scrapers/scraper.rb', line 91

def author = @declared.text(:author) || @schema.author
# (see Models::Recipe#site_name)

#canonical_url ⇒ String

Returns the canonical link of the page resolved against Models::Recipe#url, or the url itself when the page has none.

Returns:

  • (String) —

    the canonical link of the page resolved against Models::Recipe#url, or the url itself when the page has none



79
80
81
82
83
84
85
86
# File 'lib/recipe_scrapers/scraper.rb', line 79

def canonical_url
  link = document.at_css('link[rel="canonical"][href]')
  return url unless link

  URI.join(url, link["href"]).to_s
rescue URI::Error
  url
end

#category ⇒ String?

Returns the course, such as "Dessert".

Returns:

  • (String, nil) —

    the course, such as "Dessert"



99
100
# File 'lib/recipe_scrapers/scraper.rb', line 99

def category = @declared.text(:category) || @schema.category
# (see Models::Recipe#cuisine)

#cook_time ⇒ Integer?

Returns the cooking time in minutes.

Returns:

  • (Integer, nil) —

    the cooking time in minutes



122
123
# File 'lib/recipe_scrapers/scraper.rb', line 122

def cook_time = Parsers::Durations.minutes(@declared.text(:cook_time)) || @schema.cook_time
# (see Models::Recipe#prep_time)

#cooking_method ⇒ String?

Returns how it is cooked, such as "Baking".

Returns:

  • (String, nil) —

    how it is cooked, such as "Baking"



103
104
# File 'lib/recipe_scrapers/scraper.rb', line 103

def cooking_method = @declared.text(:cooking_method) || @schema.cooking_method
# (see Models::Recipe#keywords)

#cuisine ⇒ String?

Returns the cuisine, such as "Italian".

Returns:

  • (String, nil) —

    the cuisine, such as "Italian"



101
102
# File 'lib/recipe_scrapers/scraper.rb', line 101

def cuisine = @declared.text(:cuisine) || @schema.cuisine
# (see Models::Recipe#cooking_method)

#description ⇒ String?

Returns a sentence or short paragraph about the recipe.

Returns:

  • (String, nil) —

    a sentence or short paragraph about the recipe



97
98
# File 'lib/recipe_scrapers/scraper.rb', line 97

def description = @declared.text(:description) || @schema.description || @open_graph.description
# (see Models::Recipe#category)

#dietary_restrictions ⇒ Array<String>?

Returns the diets the recipe suits, such as "VeganDiet".

Returns:

  • (Array<String>, nil) —

    the diets the recipe suits, such as "VeganDiet"



111
112
# File 'lib/recipe_scrapers/scraper.rb', line 111

def dietary_restrictions = @schema.dietary_restrictions
# (see Models::Recipe#ratings)

#equipment ⇒ Array<String>?

Returns the tools a site declares, schema.org has none.

Returns:

  • (Array<String>, nil) —

    the tools a site declares, schema.org has none



107
108
# File 'lib/recipe_scrapers/scraper.rb', line 107

def equipment = @declared.rows(:equipment)
# (see Models::Recipe#nutrients)

#host ⇒ String

Returns the host of Models::Recipe#url, without "www.".

Returns:



74
75
76
# File 'lib/recipe_scrapers/scraper.rb', line 74

def host
  URI.parse(url).host.to_s.delete_prefix("www.")
end

#image ⇒ String?

Returns the address of the main image, resolved against Models::Recipe#url.

Returns:



148
149
150
151
152
153
# File 'lib/recipe_scrapers/scraper.rb', line 148

def image
  address = @declared.text(:image) || @schema.image || @open_graph.image
  address && URI.join(url, address[EMBEDDED_ADDRESS, :address] || address).to_s
rescue URI::Error
  address
end

#ingredient_groups ⇒ Array<Models::IngredientGroup>?

The ingredients in the groups the page shows, such as "For the sauce". The groups come from the HTML when the page marks them up, otherwise from the heading lines of the ingredient list. A page without groups gives one group with a nil purpose.

Returns:



138
139
140
141
142
143
144
145
# File 'lib/recipe_scrapers/scraper.rb', line 138

def ingredient_groups
  lines = ingredients
  return nil if lines.nil?

  marked = marked_groups(lines)
  groups = marked.any?(&:purpose) ? marked : Sources::IngredientGroups.from_sections(ingredient_sections)
  with_parsed_ingredients(groups)
end

#ingredients ⇒ Array<String>?

The ingredient lines as the page writes them. A line with no letter and no digit, such as "*", is a separator and left out. A line that ends in a colon and has no digit, such as "For the sauce:", is a group heading and goes to Models::Recipe#ingredient_groups instead. See Models::Recipe#parsed_ingredients.

Returns:

  • (Array<String>, nil)


127
128
129
130
# File 'lib/recipe_scrapers/scraper.rb', line 127

def ingredients
  lines = ingredient_sections.flat_map(&:last)
  lines.empty? ? nil : lines
end

#instructions ⇒ String?

Returns the steps joined with newlines.

Returns:

  • (String, nil) —

    the steps joined with newlines



135
# File 'lib/recipe_scrapers/scraper.rb', line 135

def instructions = instructions_list&.join("\n")

#instructions_list ⇒ Array<String>?

Returns the steps, with section headings as their own entries.

Returns:

  • (Array<String>, nil) —

    the steps, with section headings as their own entries



133
134
# File 'lib/recipe_scrapers/scraper.rb', line 133

def instructions_list = @declared.rows(:instructions) || @schema.instructions_list
# (see Models::Recipe#instructions)

#keywords ⇒ Array<String>?

Returns the keywords, without facet entries such as "diet: vegan".

Returns:

  • (Array<String>, nil) —

    the keywords, without facet entries such as "diet: vegan"



105
106
# File 'lib/recipe_scrapers/scraper.rb', line 105

def keywords = @declared.rows(:keywords) || @schema.keywords
# (see Models::Recipe#equipment)

#language ⇒ String?

Returns the language tag of the recipe, such as "en-US".

Returns:

  • (String, nil) —

    the language tag of the recipe, such as "en-US"



95
96
# File 'lib/recipe_scrapers/scraper.rb', line 95

def language = @declared.text(:language) || @schema.language || document_language
# (see Models::Recipe#description)

Returns the href of every link on the page, as written.

Returns:

  • (Array<String>) —

    the href of every link on the page, as written



156
157
158
# File 'lib/recipe_scrapers/scraper.rb', line 156

def links
  document.css("a[href]").map { |anchor| anchor["href"] }
end

#nutrients ⇒ Hash{String => String}?

The nutrition facts as the page publishes them, keyed by the schema.org property name. Nothing is parsed here. See Models::Recipe#parsed_nutrients.

Examples:

recipe.nutrients # => { "calories" => "219 kcal", "fatContent" => "7 g" }

Returns:

  • (Hash{String => String}, nil)


109
110
# File 'lib/recipe_scrapers/scraper.rb', line 109

def nutrients = @schema.nutrients
# (see Models::Recipe#dietary_restrictions)

#prep_time ⇒ Integer?

Returns the preparation time in minutes.

Returns:

  • (Integer, nil) —

    the preparation time in minutes



124
# File 'lib/recipe_scrapers/scraper.rb', line 124

def prep_time = Parsers::Durations.minutes(@declared.text(:prep_time)) || @schema.prep_time

#ratings ⇒ Float?

Returns the average rating.

Returns:

  • (Float, nil) —

    the average rating



113
114
# File 'lib/recipe_scrapers/scraper.rb', line 113

def ratings = @schema.ratings
# (see Models::Recipe#ratings_count)

#ratings_count ⇒ Integer?

Returns how many ratings the average is made of.

Returns:

  • (Integer, nil) —

    how many ratings the average is made of



115
# File 'lib/recipe_scrapers/scraper.rb', line 115

def ratings_count = @schema.ratings_count

#recipe? ⇒ Boolean

This method is part of a private API. You should avoid using this method if possible, as it may be removed or be changed in the future.

Returns whether the page has a title and ingredients.

Returns:

  • (Boolean) —

    whether the page has a title and ingredients



179
180
181
# File 'lib/recipe_scrapers/scraper.rb', line 179

def recipe?
  !title.nil? && !ingredients.nil?
end

#site_name ⇒ String?

Returns the name of the website.

Returns:

  • (String, nil) —

    the name of the website



93
94
# File 'lib/recipe_scrapers/scraper.rb', line 93

def site_name = @declared.text(:site_name) || @schema.site_name || @open_graph.site_name
# (see Models::Recipe#language)

#title ⇒ String?

Returns the name of the recipe.

Returns:

  • (String, nil) —

    the name of the recipe



89
90
# File 'lib/recipe_scrapers/scraper.rb', line 89

def title = @declared.text(:title) || @schema.title || @open_graph.title
# (see Models::Recipe#author)

#to_h ⇒ Hash{Symbol => Object}

This method is part of a private API. You should avoid using this method if possible, as it may be removed or be changed in the future.

Returns every field of CONTRACT with its value.

Returns:

  • (Hash{Symbol => Object}) —

    every field of CONTRACT with its value



163
164
165
# File 'lib/recipe_scrapers/scraper.rb', line 163

def to_h
  CONTRACT.to_h { |field| [field, public_send(field)] }
end

#to_recipe ⇒ Models::Recipe

This method is part of a private API. You should avoid using this method if possible, as it may be removed or be changed in the future.

Reads every field once and returns them without the page.

Returns:



172
173
174
# File 'lib/recipe_scrapers/scraper.rb', line 172

def to_recipe
  Models::Recipe.new(url: url, **to_h)
end

#total_time ⇒ Integer?

Returns the total time in minutes.

Returns:

  • (Integer, nil) —

    the total time in minutes



120
121
# File 'lib/recipe_scrapers/scraper.rb', line 120

def total_time = Parsers::Durations.minutes(@declared.text(:total_time)) || @schema.total_time
# (see Models::Recipe#cook_time)

#yields ⇒ String?

Returns how much it makes, as "4 servings" or "12 items".

Returns:

  • (String, nil) —

    how much it makes, as "4 servings" or "12 items"



118
119
# File 'lib/recipe_scrapers/scraper.rb', line 118

def yields = Parsers::Yields.parse(@declared.text(:yields)) || @schema.yields
# (see Models::Recipe#total_time)