Module: RecipeScrapers
- Defined in:
- lib/recipe_scrapers.rb,
lib/recipe_scrapers/text.rb,
lib/recipe_scrapers/errors.rb,
lib/recipe_scrapers/scraper.rb,
lib/recipe_scrapers/version.rb,
lib/recipe_scrapers/registry.rb,
lib/recipe_scrapers/site_path.rb,
lib/recipe_scrapers/declaration.rb,
lib/recipe_scrapers/http/adapter.rb,
lib/recipe_scrapers/configuration.rb,
lib/recipe_scrapers/error_tracker.rb,
lib/recipe_scrapers/http/encoding.rb,
lib/recipe_scrapers/models/recipe.rb,
lib/recipe_scrapers/parsers/chain.rb,
lib/recipe_scrapers/sites/com/app.rb,
lib/recipe_scrapers/parsers/yields.rb,
lib/recipe_scrapers/http/body_limit.rb,
lib/recipe_scrapers/models/nutrient.rb,
lib/recipe_scrapers/parsers/ratings.rb,
lib/recipe_scrapers/sites/fr/madame.rb,
lib/recipe_scrapers/sources/json_ld.rb,
lib/recipe_scrapers/parsers/quantity.rb,
lib/recipe_scrapers/sources/declared.rb,
lib/recipe_scrapers/models/ingredient.rb,
lib/recipe_scrapers/parsers/durations.rb,
lib/recipe_scrapers/parsers/nutrients.rb,
lib/recipe_scrapers/sites/de/chefkoch.rb,
lib/recipe_scrapers/sources/microdata.rb,
lib/recipe_scrapers/http/address_guard.rb,
lib/recipe_scrapers/parsers/vocabulary.rb,
lib/recipe_scrapers/sources/open_graph.rb,
lib/recipe_scrapers/sources/schema_org.rb,
lib/recipe_scrapers/parsers/ingredients.rb,
lib/recipe_scrapers/http/follow_redirects.rb,
lib/recipe_scrapers/scraper/parsed_fields.rb,
lib/recipe_scrapers/models/ingredient_group.rb,
lib/recipe_scrapers/sources/ingredient_groups.rb,
lib/recipe_scrapers/sources/schema_org/step_labels.rb,
lib/recipe_scrapers/sources/schema_org/ingredient_list.rb,
lib/recipe_scrapers/sources/schema_org/nutrition_facts.rb
Overview
Reads structured recipes from cooking websites.
The gem reads the schema.org JSON-LD or microdata a page publishes, falls back to OpenGraph, and takes a site declaration where a site publishes neither. RecipeScrapers.scrape fetches a page, RecipeScrapers.parse reads HTML you already have, and both return a Models::Recipe.
Defined Under Namespace
Modules: ErrorTracker, Http, Models, Parsers, Registry, SitePath, Sites, Sources, Text Classes: BlockedAddress, Configuration, Declaration, Error, RecipeNotFound, ResponseTooLarge, Scraper, TooManyRedirects, UnsupportedSite
Constant Summary collapse
- VERSION =
The version of the gem.
"0.2.0"
Class Method Summary collapse
-
.config ⇒ Configuration
The settings every fetch and parse uses.
-
.configure {|config| ... } ⇒ Configuration
Changes the settings in a block.
-
.parse(html, url:, supported_only: true) ⇒ Models::Recipe
Reads a recipe from HTML that is already fetched.
-
.register ⇒ Declaration
Registers a site by declaration.
-
.reset_config! ⇒ void
Throws the settings away, so the next RecipeScrapers.config call starts from the defaults.
-
.scrape(url, connection: nil, supported_only: true) ⇒ Models::Recipe
Fetches a page and reads its recipe.
Class Method Details
.config ⇒ Configuration
The settings every fetch and parse uses.
104 105 106 |
# File 'lib/recipe_scrapers.rb', line 104 def config @config ||= Configuration.new end |
.configure {|config| ... } ⇒ Configuration
Changes the settings in a block.
116 117 118 119 |
# File 'lib/recipe_scrapers.rb', line 116 def configure yield(config) config end |
.parse(html, url:, supported_only: true) ⇒ Models::Recipe
Reads a recipe from HTML that is already fetched.
61 62 63 64 65 66 67 68 69 |
# File 'lib/recipe_scrapers.rb', line 61 def parse(html, url:, supported_only: true) entry = Registry.for(url) raise UnsupportedSite, "no scraper registered for #{url}" if entry.nil? && supported_only scraper = build_scraper(entry, html, url) raise RecipeNotFound, "no recipe found at #{url}" unless scraper.recipe? scraper.to_recipe end |
.register ⇒ Declaration
Registers a site by declaration. A shortcut for RecipeScrapers::Registry.register.
97 98 99 |
# File 'lib/recipe_scrapers.rb', line 97 def register(...) Registry.register(...) end |
.reset_config! ⇒ void
This method returns an undefined value.
Throws the settings away, so the next config call starts from the defaults.
124 125 126 |
# File 'lib/recipe_scrapers.rb', line 124 def reset_config! @config = nil end |
.scrape(url, connection: nil, supported_only: true) ⇒ Models::Recipe
Fetches a page and reads its recipe.
The page is fetched through RecipeScrapers::Configuration#connection, which follows redirects, refuses private addresses and caps the body size, unless you pass your own connection.
84 85 86 87 |
# File 'lib/recipe_scrapers.rb', line 84 def scrape(url, connection: nil, supported_only: true) response = (connection || config.connection).get(url) parse(response.body, url: url, supported_only: supported_only) end |