Module: RecipeScrapers

Defined in:
lib/recipe_scrapers.rb,
lib/recipe_scrapers/text.rb,
lib/recipe_scrapers/errors.rb,
lib/recipe_scrapers/scraper.rb,
lib/recipe_scrapers/version.rb,
lib/recipe_scrapers/registry.rb,
lib/recipe_scrapers/site_path.rb,
lib/recipe_scrapers/declaration.rb,
lib/recipe_scrapers/http/adapter.rb,
lib/recipe_scrapers/configuration.rb,
lib/recipe_scrapers/error_tracker.rb,
lib/recipe_scrapers/http/encoding.rb,
lib/recipe_scrapers/models/recipe.rb,
lib/recipe_scrapers/parsers/chain.rb,
lib/recipe_scrapers/sites/com/app.rb,
lib/recipe_scrapers/parsers/yields.rb,
lib/recipe_scrapers/http/body_limit.rb,
lib/recipe_scrapers/models/nutrient.rb,
lib/recipe_scrapers/parsers/ratings.rb,
lib/recipe_scrapers/sites/fr/madame.rb,
lib/recipe_scrapers/sources/json_ld.rb,
lib/recipe_scrapers/parsers/quantity.rb,
lib/recipe_scrapers/sources/declared.rb,
lib/recipe_scrapers/models/ingredient.rb,
lib/recipe_scrapers/parsers/durations.rb,
lib/recipe_scrapers/parsers/nutrients.rb,
lib/recipe_scrapers/sites/de/chefkoch.rb,
lib/recipe_scrapers/sources/microdata.rb,
lib/recipe_scrapers/http/address_guard.rb,
lib/recipe_scrapers/parsers/vocabulary.rb,
lib/recipe_scrapers/sources/open_graph.rb,
lib/recipe_scrapers/sources/schema_org.rb,
lib/recipe_scrapers/parsers/ingredients.rb,
lib/recipe_scrapers/http/follow_redirects.rb,
lib/recipe_scrapers/scraper/parsed_fields.rb,
lib/recipe_scrapers/models/ingredient_group.rb,
lib/recipe_scrapers/sources/ingredient_groups.rb,
lib/recipe_scrapers/sources/schema_org/step_labels.rb,
lib/recipe_scrapers/sources/schema_org/ingredient_list.rb,
lib/recipe_scrapers/sources/schema_org/nutrition_facts.rb

Overview

Reads structured recipes from cooking websites.

The gem reads the schema.org JSON-LD or microdata a page publishes, falls back to OpenGraph, and takes a site declaration where a site publishes neither. RecipeScrapers.scrape fetches a page, RecipeScrapers.parse reads HTML you already have, and both return a Models::Recipe.

Examples:

Fetch and read a recipe

recipe = RecipeScrapers.scrape("https://www.recipetineats.com/crispy-potato-straws-pommes-paille/")
recipe.title       # => "Crispy potato straws (Pommes Paille)"
recipe.ingredients # => ["1 potato (Aus: Sebago, US: russet, UK: Maris Piper), ...", ...]

Defined Under Namespace

Modules: ErrorTracker, Http, Models, Parsers, Registry, SitePath, Sites, Sources, Text Classes: BlockedAddress, Configuration, Declaration, Error, RecipeNotFound, ResponseTooLarge, Scraper, TooManyRedirects, UnsupportedSite

Constant Summary collapse

VERSION =

The version of the gem.

"0.2.0"

Class Method Summary collapse

Class Method Details

.config ⇒ Configuration

The settings every fetch and parse uses.

Returns:



104
105
106
# File 'lib/recipe_scrapers.rb', line 104

def config
  @config ||= Configuration.new
end

.configure {|config| ... } ⇒ Configuration

Changes the settings in a block.

Examples:

RecipeScrapers.configure do |config|
  config.timeout = 30
end

Yield Parameters:

Returns:



116
117
118
119
# File 'lib/recipe_scrapers.rb', line 116

def configure
  yield(config)
  config
end

.parse(html, url:, supported_only: true) ⇒ Models::Recipe

Reads a recipe from HTML that is already fetched.

Parameters:

  • html (String) —

    the page source

  • url (String) —

    the address of the page, used to pick the site and to resolve relative links

  • supported_only (Boolean) (defaults to: true) —

    raise for a host the gem does not support. Pass false to read any page that publishes schema.org or OpenGraph markup

Returns:

Raises:

  • (UnsupportedSite) —

    when the host is not registered and supported_only is true

  • (RecipeNotFound) —

    when the page has no title or no ingredients



61
62
63
64
65
66
67
68
69
# File 'lib/recipe_scrapers.rb', line 61

def parse(html, url:, supported_only: true)
  entry = Registry.for(url)
  raise UnsupportedSite, "no scraper registered for #{url}" if entry.nil? && supported_only

  scraper = build_scraper(entry, html, url)
  raise RecipeNotFound, "no recipe found at #{url}" unless scraper.recipe?

  scraper.to_recipe
end

.register ⇒ Declaration

Registers a site by declaration. A shortcut for RecipeScrapers::Registry.register.

Examples:

RecipeScrapers.register "example.com" do
  title "h1.recipe-title"
  ingredients rows: "ul.ingredients li"
end

Returns:



97
98
99
# File 'lib/recipe_scrapers.rb', line 97

def register(...)
  Registry.register(...)
end

.reset_config! ⇒ void

This method returns an undefined value.

Throws the settings away, so the next config call starts from the defaults.



124
125
126
# File 'lib/recipe_scrapers.rb', line 124

def reset_config!
  @config = nil
end

.scrape(url, connection: nil, supported_only: true) ⇒ Models::Recipe

Fetches a page and reads its recipe.

The page is fetched through RecipeScrapers::Configuration#connection, which follows redirects, refuses private addresses and caps the body size, unless you pass your own connection.

Parameters:

  • url (String) —

    the address of the recipe page

  • connection (Faraday::Connection, nil) (defaults to: nil) —

    the connection to fetch with

  • supported_only (Boolean) (defaults to: true) —

    see parse

Returns:

Raises:



84
85
86
87
# File 'lib/recipe_scrapers.rb', line 84

def scrape(url, connection: nil, supported_only: true)
  response = (connection || config.connection).get(url)
  parse(response.body, url: url, supported_only: supported_only)
end