[pull] main from freeCodeCamp:main - #275
Merged
Merged
Conversation
Add ruby/rack docs
Fixes #2715.
…ockfile Update dependency image_optim to v0.32.0
Symlink into ~/.config/fish to use:
ln -s $PWD/fish/functions/*.fish ~/.config/fish/functions/
ln -s $PWD/fish/completions/*.fish ~/.config/fish/completions/
Run the thor tasks in the current directory instead of looking up the checkout root.
Alias for the maintainer workflow, mirroring `npm outdated` and `bundle outdated`. Also completed by the fish function.
Scraping rewrites every page file, so the default size and mod-time comparison re-uploads all files of a documentation even when only a few pages changed.
A documentation consists of thousands of small files, whose upload is bound by request round-trips: rclone defaults to 4 parallel transfers, the AWS CLI to 10 concurrent requests.
Add pytest documentation (9.1.1)
The entries filter kept a process-wide list of entry names and returned no entries for any page whose name had been seen before. Since a page without entries is not stored at all, this silently dropped whole documents rather than just index entries. Unrelated pages legitimately share a name: every container's erase_if page is titled "std::erase_if (std::<container>)", which normalises to plain "std::erase_if", so only the first one crawled survived. Same for std::move in <algorithm> vs <utility>, and std::optional::operator bool. The deduplication exists because cppreference serves the same wiki page under several URLs and the scraper stores pages under the requested URL rather than the effective one. Key it on the canonical page name from the "Retrieved from" footer instead, which identifies the underlying wiki page exactly. Fixes #2175 Fixes #2190 Fixes #2223
UrlScraper now stores every response it fetches in tmp/cache/<slug> and serves subsequent runs from there, which makes iterating on a scraper's filters a lot faster. Only successful responses are stored, so transient failures aren't pinned forever. Cached responses are collected by the Requester and handed over iteratively rather than from within Hydra#add, because delivering them right away would nest one request's callbacks inside the previous one's and overflow the stack on large documentations. thor docs:clean deletes the caches. It recognizes them by a marker file so that it leaves the assets cache alone.
The cache files were Marshal dumps, which are opaque when you open one to find out what a scraper actually got back. They're now JSON, in the entry schema of the HTTP Archive format: http://www.softwareishard.com/blog/har-12-spec/ An archive is a log of many entries; keeping one entry per file instead means the cache stays incremental, at the cost of the files not being valid archives on their own. Bodies that aren't valid UTF-8 fall back to the spec's base64 encoding, and are handed back to the scrapers as binary either way, so a warm run sees exactly what a cold one does. Cache files gained a .json extension, so the entries written by the previous format are ignored; run thor docs:clean to drop them.
Cache scraper responses in tmp/cache
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.4)
Can you help keep this open source service alive? 💖 Please sponsor : )