At a glance: which path for which site
Rowvana handles three kinds of job. Work out which one your site is first, and the rest is easy.
A. A plain list or table
Many similar items on one page: a product grid, search results, an HTML table, a news list.
→ Pick elements or Auto-detectB. A shop where each product has its own page
The list shows only name and price; everything else (SKU, seller, specs) is on the product page.
→ Product detail pages (link mode)C. Maps / directories
Clicking an item does not open a new page; a panel opens beside the list on the same page, like Google Maps or Yelp.
→ ⚡ Auto-setup (one button)The key idea: what you see in your browser is what gets scraped. Pages behind a login and content loaded by JavaScript all work. Your data goes to no server; everything happens in your browser.
🤖 robots.txt: on by default
You will find the "Respect robots.txt" card in the ⚙ tab. If a page is not allowed, Setup shows a red line and a ⚙ Details button. Rowvana reads each site's robots.txt file, where the site lists the paths automated clients should not visit.
- When allowed, a green line appears. A
Crawl-delayis respected too: if it is longer than your delay, the site's value is used. - When not allowed, the run does not start, and you are shown the line that blocks it (
Disallow: /private/). - Each product page is checked separately. If the list is allowed but the product path is not, that row says why and the rest of the job carries on.
You can turn the switch off, for example when a client has given you written permission for their own site that robots.txt cannot express. But remember: robots.txt is not a licence. Being allowed to crawl does not give you the right to republish data; the site's Terms of Service apply separately.
🛰 Using the site's own API
Most modern sites load their list from a JSON endpoint. That JSON often holds fields the page does not show: latitude and longitude, email, opening hours, internal IDs. Taking the data from there is more accurate, richer and much faster than reading the HTML, and pagination needs no clicking, only a number that changes.
- In the Setup tab, turn on 🛰 Use the site's API.
- Press Find the API. The page reloads once, so calls made while loading are caught too.
- Search or scroll the site normally, the way a person would.
- Press Check again. You get a list of the JSON that arrived, with the richest one already selected.
- Check the sample columns below → ▶ Extract from the API.
- It finds where the records are by itself:
data.dealers,results.hits, or the top level. If there are several candidates, pick one from the dropdown. - Nested keys become columns:
contact.phone,address.city. Arrays of plain values are joined (Bosch | Cube). - Paging is detected: it finds
page, oroffset+limit(in the URL or the POST body), and calls the endpoint again. The From/To pages in the Pagination card apply here too. - The request is replayed exactly, with its headers (including auth tokens) and your own cookies, so it works on sites you are logged in to.
How it works: Chrome does not let extensions read response bodies (webRequest sees only headers). So a hook is placed on the page's own fetch and XMLHttpRequest. It only watches; it changes nothing and sends no requests of its own. Turning the switch off removes the hook. robots.txt applies here as well.
Sites that search with POST: many store finders send the location and radius in a POST body. That body is captured in every form (string, FormData, URLSearchParams, or inside a Request object), and the response already captured is used as the first page, so no extra call is needed. If something comes back empty, the log says why.
How the right API is chosen
On sites like Rightmove a single search page sends more than 300 requests, most of them JSON for ads, consent and analytics. To find the real listing API, Rowvana does three things:
- Ads and trackers are never captured (prebid, bid, cookie_sync, vendorlist, consent…). When space runs out, older calls from other sites are dropped instead of new ones, so a listing call that arrives late is not lost.
- The choice is made by matching the page, not by counting records. Addresses, names and prices from each JSON's records are checked against what is visible on the page, along with whether it is the same site, whether the URL carries the page's search settings, and whether it has pagination info. The list shows the reason next to each one: "✓ 25 of 25 records are on this page".
- Paging comes from the API's own answer. Patterns like
index=0, 24, 48…are recognised, the step is taken from the API'snext/pageSize(not by counting records), and "From 1 To 3" means the real first three pages of results.
Fully automatic
- The yardstick is the page's main list. Rowvana finds the largest, most informative repeating list on the page (for example property cards) and checks whether each API record matches one card in it. Link text such as "Central London" in a sidebar is not a card, so it is no longer chosen.
- It presses Next by itself. On many sites (Rightmove included) the first page comes with the HTML, so after a reload there is no API call at all. When you press Find the API and no list is found, Rowvana presses Next once (or Load more, or scrolls), the site calls its API, and that call is captured and checked. The log says what was pressed. Check again only looks; it presses nothing.
If nothing matches even then, Rowvana does not guess. It tells you to change the page, sort or filter on the site and press Check again. If that still fails, the site builds its list in HTML, so use 🎯 Pick instead.
Not every site has an API. On older server-rendered pages everything arrives as HTML, so "Find the API" finds nothing; use the 🎯 picker as usual. WebSocket and GraphQL subscriptions are not captured yet.
Install
- Open Rowvana on the Chrome Web Store (or press Add to Chrome on rowvana.com).
- Press Add to Chrome and confirm.
- Click the puzzle icon in the toolbar and pin the ⚡ Rowvana icon, so it is always in reach.
- Open the popup and Sign in with Google. The Free plan costs nothing.
Updates install themselves: Chrome checks for a new version every few hours. You do not need to do anything.
Your first scrape in 5 minutes
Start with a simple page such as books.toscrape.com (a site made for scraping practice). The steps below are the same as in the video.
Path 1: Auto-detect (fastest)
- Open the page and click the ⚡ icon → the popup opens.
- Press ✨ Auto-detect. Every repeating structure on the page (product grids, tables) is listed, with its number of rows and columns and a sample.
- Click the one you want. The row container and fields fill in by themselves.
- Press Preview. It scrapes only this page and shows it in the Results tab. Check it looks right.
- If it does, press ▶ Extract data. When it finishes, go to Results → CSV / Excel.
Path 2: Pick elements (choose for yourself)
If you don't like what Auto-detect found, or you only need a few things:
- In the popup press 🎯 Pick elements on page. The popup closes and a small panel appears in a corner of the page.
- Click any one card or item (a book, a product). All the other cards get a green dashed outline, meaning the row container was found. The panel moves on to the "Fields" step by itself.
- Now click, one by one, the things you want inside that card: name, price, image. Each click is one column, marked in orange on every card.
- In the panel you can rename a column and choose what to take:
text,href(link) orsrc(image URL). - Save & Close → open the popup → ▶ Extract data.
Clicked something by mistake? Press ✕ next to that field in the panel. To cancel everything, press Esc or Cancel.
Path 3: 🎯 and ✨ next to every box
You don't have to open the big picker panel. Wherever you would type a selector or XPath (the row container, each field, each product-page field, and the Advanced boxes for detail panel / click / close / specs / "open this") there are two buttons beside it:
- 🎯: click the thing on the page and its selector goes straight into that box. Hovering shows what the selector will be and how many rows it matches.
- ✨: finds it for you. For fields it goes by the name: "Price" → the price, "Image" → the image (src), "Link" → the link (href), "Title", "Rating", "Address", "Phone", "Date"… If the name is unfamiliar it looks for that word in the page's classes. For names like "Field 7" it asks you to point with 🎯.
CSS or XPath: both really work
Choosing XPath in the dropdown converts the selector itself to XPath. The conversion uses classes, not positions like div[3], and is checked on the page to make sure it still matches the same things. 🎯, ✨ and the big picker all follow the box's type. Where an exact conversion is not possible, the dropdown switches back and says why.
Path 4: from the JSON in the page source
On sites like Rightmove, view-source shows window.PAGE_MODEL = {…} or __NEXT_DATA__. That is where the real data lives. Page class names are random and change often, so reading from here is the most durable option.
- Choose JSON in the field's dropdown. If the field already has a CSS selector, Rowvana finds where that text sits in the data and fills in the path, for example
PAGE_MODEL.propertyData.prices.primaryPrice. - On an empty JSON field, ✨ searches by the field's name (Price, Bedrooms, Address, Image…), and 🎯 on some text in the page gives you its path.
[*]means all of them:PAGE_MODEL.propertyData.images[*].urlgives every image link, and with 🖼 on every image is downloaded.- On a list page, 🛰 Find the API shows
📄 __NEXT_DATA__ (in the page)in its list. Choose it and rows come from each page's JSON. Keep Next on in Pagination.
If the row container picks up footer links by mistake (Title/Price empty, Link only "stamp-duty-calculator"), the run stops and nothing is exported. You are told which fields are empty and how to fix it.
The picker panel: what the colours mean
Each highlight colour you see on the page means one thing:
The panel's buttons
- 1. Row / Card · 2. Fields
- Which mode you are in. Row first, then fields. You can switch by hand.
- ✨ Auto-detect rows
- Tries to do the whole job: rows and likely fields together.
- Clear
- Removes the row container; fields stay.
- Save & Close
- Sends everything to the popup and closes the panel. Open the popup again and everything is filled in.
- – (top right)
- Shrinks the panel, handy when it covers the page. You can also drag it by its header.
If you click fields without a row container, Rowvana uses "column mode": each field takes everything it finds on the whole page as a column, and the columns are joined in order. If one field matches fewer times, rows can shift. So for a clean table, set the row container first.
The popup comes back by itself after Save: the popup has to close for the picker to work (Chrome closes toolbar popups when the page gets focus). Rowvana then reopens the popup for you, with the picked fields filled in. On older Chrome versions where that is not possible, the toolbar icon shows a green badge with a number and the page says "Saved 5 field(s) — click the ⚡ toolbar icon to carry on."
Every part of the popup
The popup has three tabs: Setup (what to scrape), Results (what you got) and Saved (save/load setups), plus ⚙ for settings. The Setup tab from top to bottom:
- Current tab
- The page the job runs on. The popup always uses the tab in front.
- 🎯 Pick elements on page
- Starts the picker (step 3).
- ✨ Auto-detect
- Finds the tables and lists on the page; pick one.
- ⚡ Auto-setup: list + detail panel
- One-button setup for Maps-style sites (step 7).
- Repeating row container
- The selector for each record, e.g.
.product-card. CSS or XPath can be chosen beside it. Test shows how many match. - Fields
- Each column: name, what to take (text/href/src/…), selector. + Add field adds one by hand. Test shows a green/red number by each field: how many rows it matched.
- Pagination
- When there are several pages (step 5).
- 🛒 Product detail pages
- Going inside each item (step 6, 7). Nothing opens unless the Allow switch is on.
- ▶ Extract data · Preview
- Extract = the full job (with pagination and details); it runs in the background and does not stop when you close the popup. Preview = this page only, for a quick look.
- Log box
- What is happening, line by line: which page, how many found, where it stopped.
A field's "what to take" options
| Option | What you get | When |
|---|---|---|
| text | The element's text | Names, prices, descriptions: most of the time |
| href | The link's full URL | Product/profile links (always an absolute URL) |
| src | The image URL | Thumbnails, product images |
| title / alt | That attribute's value | Tooltips, image descriptions |
| value | An input's value | Form fields |
| innerHTML | The inner HTML | When you need the formatting |
| custom | Any attribute | e.g. data-id, data-price; type the name in the box below |
Pagination: more than one page
Tick Enable in the Pagination card, then choose the mode that matches how the site changes pages:
| Mode | When | What to give |
|---|---|---|
| Next button / link | There is a "Next »" button or link at the bottom | That button's selector (e.g. li.next a, a[rel=next]) |
| URL pattern | The URL has the page number: ?page=2, /page/3/ | The pattern with {page} in place of the number: https://site.com/list?page={page} |
| Infinite scroll | More items appear as you scroll down | Max scrolls; the "Load more" button's selector if there is one |
Don't type the Next selector by hand. Press ✨ Find it, or use 🎯 Pick "Next" on page and click the button. Both give a selector that works on every page (e.g. a.s-pagination-next, a[aria-label^="Go to next page"]).
Don't use DevTools' "Copy selector". It gives something like #search > div > … > li:nth-child(8) > span > a, an address in the page layout that changes on the next page (on Amazon by page 3). Labels like "Go to next page, page 2" have the same problem: they match only page 1. The popup warns you about selectors like these, and if a selector stops matching in the middle of a run, Rowvana finds the real Next by itself and says so in the log.
With infinite scroll, the row container is one item, not the whole list. If you take the box around the whole list, it matches only one thing on the page and new items arriving on scroll go unnoticed. Rowvana detects and fixes this (the log says "matches only 1 element — it is the box around the list"). Feeds that keep only one screen of items in the page are read while scrolling, so items at the top are not lost. On slow sites, set "Delay between pages" to 3000–5000 ms.
The site's own limits: even if you set To page 1000, you get no more than the site shows. Airbnb shows 15 pages per search, Amazon about 20. The log says "You asked for up to page 1000, but the site shows no Next link here". To get more, split the search (price ranges, areas, filters).
When product pages are on too: check Max items (0 = all) in the Product Page card. The default is 0, meaning all of them.
Two safety nets are built in: the same row arriving twice is dropped, and the crawl stops by itself when a page brings nothing new, so a high page limit does no harm. Don't set Delay between pages below 700 ms; going too fast can get you blocked.
Only 3 of 10 pages: page range
The Pagination card has From page and To page, and that is the only control. 1 to 3 means the first three pages; 4 to 6 means exactly those three. (In infinite scroll mode, Max scrolls takes its place.)
- In URL pattern mode those pages are opened directly; earlier ones are never visited, so it is fast.
- In Next button mode there is no other way to reach page 4, so pages 1–3 are clicked through but not scraped. The log shows "Page 1: skipped" and the status "Skipping to 4".
- Each row's
_pagecolumn holds the real page number (4, 5, 6). - Infinite scroll has no pages, so the option isn't shown there; Max scrolls controls it instead.
This is very handy for big jobs: give a client a 2-page sample first (1–2), and if they like it, take the rest range by range (3–10, 11–20). Each run's file is separate, so merge them at the end. Saved profiles keep the range too.
No selectors to type
You don't need to write CSS in the "Next-page selector" box. The Pagination card has two buttons:
- 🎯 Pick “Next” on page: the popup closes and a highlighter starts on the page. Hovering shows which selector will be used; click the Next button and it goes into the box. The click does not reach the page, so you won't move to the next page by mistake. Esc cancels.
- ✨ Find it: you press nothing; it searches by itself (
rel="next",aria-label, "Next / › / »" text, pagination containers), picks the most likely one and writes below why it chose it.
Both build selectors that also work on later pages (a[rel=next], aria-label or a class) rather than a long DOM path, because the Next button's position changes from page to page.
In infinite scroll mode the same two buttons are there for a "Load more" button. Things like Add to cart / Checkout / Sign in are never suggested.
Shops: opening every product page for full details
A shop's list page shows only name, price and rating. Brand, SKU, seller, warranty and specifications are on each product's own page. In this mode Rowvana scrapes the list, then visits each product link and adds the extra details to the same row.
- On the list page, pick the row and fields as before. One field must be the product link: click the product name and choose
hrefin the dropdown. (For both name and link, click the same spot twice and keep one as text, one as href.) - In the popup, turn on Allow in the 🛒 Product detail pages card. "How the details open" stays at Follow a link to a product page.
- In Product link column, choose the column with the href (it says "— link" after the name).
- Now open any one product page (any one will do; they all share a template). Open the popup and press 🎯 Pick fields on this product page. The picker appears: click the price, seller, brand, rating. Save & Close.
- Go back to the list page and press ▶ Extract list + product pages.
The result arrives in one CSV like this:
Title, Price, Link, Seller, Was, Brand, _page, _source_url
How it runs: product pages load one after another in a single hidden background tab; your own tab stays where it was, and the background tab closes itself at the end. A link that appears twice is opened once. If a page fails to open, that row's _detail_error column says why, and the rest carries on.
The 🛒 card has Max items (0 = all), and ⚙ → 🛒 Product pages has Delay per item (default 1000 ms). 200 products × 1 second = 3–5 minutes. Don't go below 500 ms.
Pick exactly the fields you want with 🎯. If one sits in a hidden section, turn on ⤢ beside it, and only that section is opened.
⤢: open just one field's section
On pages like Airbnb, information is hidden in separate sections: the Description has its own "Show more", Facilities its own "Show all 31 amenities", and House rules is a closed <details>. So every product-page field has a small ⤢ button beside it.
With ⤢ on, only that field's own section is opened before reading it; the rest of the page is left alone. The control is searched for in this order:
- An
aria-controlspointing to the id of the field's container (well-built accordions work this way) - The
<summary>of a closed<details>above the field - The nearest "Show more / Show all / View full…" button in the section
If the field isn't on the page yet (inside a closed accordion), the selector is shortened from the end (if #rev-body p doesn't match, #rev-body) so the button can be found through its container.
With ⤢ on, a small box appears below: leave it empty for auto. If the wrong button gets pressed, type your own selector there (e.g. #amenityToggle) and exactly that is pressed. When a pop-up opens, its value is taken before it is closed, because that is the only chance.
The safety rules are the same: even if you type #addToCart yourself, it won't be pressed.
If you pick a field again, the ⤢ setting of a field with the same name is kept; you don't have to turn it on again.
⤢ on list columns too: each row's own "Show more"
On list pages each card can have its own "Show more". So the FIELDS card in Setup has a ⤢ button beside each column too. With it on, each row's own button is pressed once before the page is read.
It is designed for speed: pressing one row and waiting before the next would take a minute for 100 rows. So all rows' buttons are pressed first, then there is one wait (default 700 ms). If a row's control opens a pop-up, it is closed, because you cannot tell which row the pop-up belongs to. At most 120 rows per page.
It never presses twice. After opening, "Show more" turns into "Show less"; pressing it again would close what was opened. Controls with aria-expanded="true", an open <details>, or text like "Show less / Hide" are skipped, and the log says "already open".
🎯 and ✨: no selectors to type
With ⤢ on, two buttons appear beside the selector box:
- 🎯 Pick: the popup closes and a highlighter starts on the page. Click the button that opens the section and the selector goes in by itself. The click does not reach the page, so nothing opens by mistake. For list fields the selector is built to work on every row (not an id, which would open only the first card).
- ✨ Find: you press nothing; it works out the control from the field's selector, fills the box and writes how it found it, e.g. "Found “Show more” by aria-controls — matches in 2 of 2 rows".
An empty box still means auto: the search happens during the run. 🎯/✨ simply show it to you before you run.
🖼 Image download: the real files, not just URLs
Having an image's address in a sheet is not the same as having the image: when a site changes its CDN, old links die. So in both sections (FIELDS and Product Page fields) every field has a 🖼 button. With it on, the image is saved as a real file in Downloads/uws-images/<site>-<date-time>/, and the sheet gets a new column, Photo (file).
Default: one image per row. This is deliberate: 100 rows × 8 gallery images is 800 files, which is annoying when you didn't ask for it. For the whole gallery, choose All images in the gallery in the dropdown; then the other images around the one you picked are saved too (the nearest box holding several images counts as the "gallery"), with no extra selector to write.
There is a Test button beside it that tells you "3 image(s) per row × 20 row(s) ≈ 60 file(s) per page", so you know how many files are coming before you run.
Getting the right URL is the real work, and it is done carefully:
- With
srcset, the highest resolution is taken - Lazy-loaded images have a 1×1 transparent placeholder in
srcand the real one indata-src; the real one is taken - Images drawn with CSS
background-imageare caught too - 24px icons and badges are skipped; tiny pictures aren't photos
- The same URL is never saved twice, and robots.txt applies here too
File names and URLs side by side
Each file is named after the image's own URL: …/villa-pool.jpg is saved as villa-pool.jpg. And the URL is kept in two places:
- Two columns side by side in the sheet, in the same order:
Photo (file)andPhoto (image url) - An
images.csvin each run's folder:file, image_url, row, field, page_url, status. Images that were not saved (robots.txt, file limit reached, server error) are listed with the reason, never lost silently.
Several images get 1, 2, 3 at the end of the name. With "All images in the gallery", by their place in the gallery: pool_1.jpg, bedroom_2.jpg, kitchen_3.jpg. If images from different rows share a name (many CDNs serve …/large.jpg for every product), the second becomes large_2.jpg, the third large_3.jpg; nothing is overwritten.
📁 Output folder: a folder of your choice
- In ⚙ → Images, press 📁 Choose output folder…. A settings page opens.
- 📁 A folder I choose → Choose folder…. Your system's folder dialog appears; choose any folder on any drive (e.g.
D:\Client Photos). - If Chrome asks for access, press Allow. From then on, all images go there.
If you'd rather avoid permissions, use the other option: ⬇ Inside the Downloads folder, giving the name of a folder inside Downloads (default uws-images).
Two things to know. (1) Chrome shows only the folder's name, not the full path; that is a browser security rule. (2) After a Chrome restart it sometimes asks for access to the folder again; press Re-allow access on the settings page. Until then the run does not stop: images go to Downloads and the log says why.
A new folder for each run, by date or domain
On the settings page, choose under Name each run’s folder by:
| Option | Folder name |
|---|---|
| Domain + date (default) | airbnb.com_2026-09-21_17-33-35 |
| Date & time | 2026-09-21_17-33-35 |
| Domain | airbnb.com → next run airbnb.com-2 → airbnb.com-3 |
If the name already exists, -2, -3 is added, so a new run never mixes with an old one. A live preview on the page shows what the result will look like.
The Max image files per run box in the FIELDS card (default 500) appears only when 🖼 is on for some field.
Maps / directories: ⚡ Auto-setup in one button
On Google Maps, clicking a result doesn't open a new page; a panel appears beside the list on the same page (address, phone, website, hours). There is no link to follow, so Rowvana clicks each item itself, reads the panel, closes it and moves on.
The easy way (enough almost every time)
- Search on Maps (e.g. "restaurant london") so the results list appears on the left. You don't need to open any result yourself.
- Open the popup and press ⚡ Auto-setup: list + detail panel.
- Wait 10–15 seconds. Rowvana finds the list → clicks the first item → recognises the panel → suggests name/address/phone/website → clicks a second item to check that different data comes back.
- When the log says "Ready — press Run", every box is filled. Look over the Fields list once and remove columns you don't need with ✕.
- ▶ Extract list + product pages. When it finishes, Results → Excel.
Pressing this button means you allow clicks on items, so the Allow switch turns on by itself. The switch is visible; turn it off if you don't want that, and only the list is scraped.
The most common mistake: pressing Auto-setup while on a single place's page (URL with /place/...). There is no results list there, so Rowvana finds only two or so fake "rows". Go to the search results page first; the list of cards must be visible on the left. If fewer than 3 rows are found, Rowvana stops and tells you this.
For more results
Keep the Maps tab in front while it scrolls. Chrome slows timers in background tabs, and Google then stops loading new results, so scrolling seems "stuck". Closing the popup is fine (the job runs in the background); just don't switch tabs.
The Maps list is virtualised: only a few cards are on the page until you scroll. Auto-setup notices this and turns on infinite scroll by itself (Max scrolls 25), and says so in the log. For more, raise it to 50–80. Rowvana scrolls the results list's own scrollbar, not the whole page, which is why it works.
Need thousands of records? Google itself returns at most ~100–120 results per search; that is not a Rowvana limit. So split the search: "restaurant in Camden", "restaurant in Soho", "restaurant in Shoreditch"… each a separate run, then combine the files. Set it up once and keep it as a saved scraper, and each search is just ▶. (Check the site's Terms of Service before scraping.)
The manual way (if Auto-setup gets it wrong)
- With 🎯 Pick elements on page, click a result card to set the row, click the name and rating for fields, then Save.
- Turn Allow on in the 🛒 card → "How the details open" = Click the item, panel opens on the same page.
- Press 🖱 Click first item & pick fields. Rowvana clicks the first item, opens the panel and starts the picker there (the panel has purple dashes).
- Press ✨ Auto-detect fields in the panel, or click the address/phone/website by hand; both together work too. Save & Close.
- Run.
Good to know
- If the panel doesn't change after a click, that row gets
_detail_error("The panel did not change after clicking…"); the previous panel's data is never copied by mistake. - Rowvana waits until the panel goes from "Loading…" to real content; it never reads a half-loaded panel.
- If the panel has a Close button, you can give its selector in ⚙ → 🛒 Product pages → "Close button"; it works without it too, because clicking the next item replaces the panel.
Results and export
Excel cannot hold more than 32,767 characters in one cell. Longer text is cut only in .xlsx, and the export message says so; CSV and JSON keep all of it. Very long numbers (such as IDs) go to Excel as text so no digits are lost.
The Results tab shows the first 200 rows in a table (exports have them all). The top shows the total rows and how many pages they came from.
- CSV
- UTF-8 with BOM, so non-English text shows correctly in Excel. File name:
sitename-date-time.csv. - JSON
- All the data, for use in programs.
- Excel
- A real
.xlsx: bold header, frozen first row, auto-filter on. Numbers come in as numbers (but text like "007" that starts with zero stays text). - Copy
- Tab-separated to the clipboard; paste straight into Google Sheets with Ctrl+V.
- Clear
- Deletes the results (the setup stays).
Column order: your fields in order → detail fields → auto specs → finally _page (which page) and _source_url. If column names clash, "(spec)" is added to the spec column's name; a column you picked is never removed.
💾 Saving a scraper
Once it is set up, don't lose it. Right under the ▶ button in the Setup tab is 💾 Save this scraper: one click, no need to type a name (one is suggested from the page title, e.g. "restaurant london - Google Maps").
What is saved: the row selector, every column, the pagination type and numbers, all product page / panel fields, and the Allow switch, meaning the whole scraper, not a half-filled form.
- Saving again for the same list on the same site shows 💾 Update "…": the old one is updated, no duplicates pile up. For a separate copy, use Save as new beside it.
- The Saved tab says under each one what it actually does:
8 columns · panel (9) · scroll. Rename with ✎; Load fills every box and takes you back to Setup, with only ▶ left to press.
For thousands of Maps records this is the real workhorse: save once, then for each new search (another area or category) just Load → ▶. No setting up again each time.
A saved scraper may not work on a different kind of page (a profile made for search results, used on a single place's page). In that case the run stops and tells you; rebuild it with ⚡ Auto-setup and press 💾 Update.
Below are some Starter templates: quotes.toscrape and books.toscrape (for practice), any HTML table, all links on a page, all images. Use fills in the setup.
⚙ Settings: proxy, API keys and Reset
At the end of the tab row there is just a ⚙ icon (to save space; hover to see its name). Everything you set once lives here: robots.txt, page load timeout, proxy, API keys, the image folder, your account and, at the very bottom, Delete account. Reset is the ↺ button at the top.
Not sure what a button or switch does? Hover over it; every control has a short explanation.
⏱ Page load timeout
Rowvana waits for each page to load fully. But one stuck page won't hold up the whole run: if it doesn't load within the time set here (default 30 seconds), that page is skipped and the run moves on. The log says which page was skipped, and a summary comes at the end.
- With page numbers in the URL: straight to the next number.
- With a Next button: it finds Next in whatever part of the page did load. If nothing loaded, there is no way forward, so it stops and says so.
- Product page: that row gets the reason, and the next product runs.
- API mode: the stuck call is skipped and the next page runs.
Increase the time on slow sites (60–90 seconds); too low and good pages get skipped too.
🛡 Proxy: avoiding blocks
Sites block scrapers by IP. Pull many pages and you may suddenly get "Access denied" or a CAPTCHA. There are two options, for two different jobs:
1) Your own proxy list
This is for normal crawls, because pages are read from a browser tab, so the tab itself has to go through the proxy. One per line, pasted the way your provider gave them:
203.0.113.5:8000
203.0.113.6:8000:username:password
username:password@203.0.113.7:8000
socks5://203.0.113.8:1080
- When the IP changes: once per run · every page · every N pages · only when blocked.
- Blocks are detected automatically. When a wall appears instead of data ("Access denied", "unusual traffic", a CAPTCHA frame), Rowvana notices, changes IP and tries again (twice), and tells you if it still fails. With "stop when blocked" set, it stops, keeping the rows collected so far.
- Proxies with a password just work; Chrome's sign-in box does not appear, Rowvana answers it.
- Test the first proxy sends one request and shows which IP the other side sees, so you know the list works before you start.
In Chrome the proxy setting applies to the whole browser, so your other tabs also use the proxy while a run is going. It is set when the run starts and removed by itself when it ends (even if it fails). To remove it by force, press Reset browser proxy.
2) A proxy API service (with an API key)
ScraperAPI, ScrapingBee, ZenRows, ScrapingAnt, or your own URL pattern ({key} and {url} are filled in). They fetch on your behalf but cannot draw the page in a tab, so they don't work for normal point-and-click crawls. They work where Rowvana fetches by itself: API mode, image download, and optionally robots.txt. Tick which ones to use them for.
Your proxy lists and keys stay on your computer; they are never sent to Rowvana.
🔑 API keys
Name + header + key. When Rowvana sends a request itself (API mode, image download), the key goes as a header, e.g. Authorization or X-API-Key. Useful for taking data from a client's own API.
Keys are saved in this browser profile unencrypted (like a saved password). Anyone with access to this computer can read them; keep that in mind for a client's keys.
↺ Reset: a clean start
Reset is the ↺ button at the top. Useful when starting work for a new client. Tick what to remove: current setup · results · saved scrapers · image settings · proxies and API keys · log · sign-in. Reset the ticked items asks you to confirm first and shows exactly what will be removed.
To clear everything at once, Reset everything removes your setup, results, scrapers, settings, proxies and keys and signs you out (other accounts' data in this browser is not touched). After any Reset you are taken to the top of the Setup tab with "Reset done."
👤 Your Rowvana account and plan
You don't need to sign in to look around. All tabs, Auto-setup, the picker, selectors, preview and hover tips work as they are. But ▶ Run, Export/Copy, Save, 🛰 Find the API, image preview/download and proxy Test show a panel: "Sign in to run this scraper. The Free plan costs nothing." After Sign in with Google, whatever you pressed runs by itself. Not now closes the panel and leaves your setup as it was.
The top line (on every tab): when signed out, Sign in and ↺. When signed in: your photo, name, email, a plan badge, ⏻ (sign out) and ↺ (reset). The badge colour shows where your plan period is: green = new, blue = midway, yellow = near the end, red = almost over, grey = Free. Hover over the badge to see when it started and when it renews (or ends).
Each account's data is separate. Setup, results, saved scrapers, settings, proxies and API keys are all kept per account. After signing out the popup looks empty, so the next person on the same Chrome sees nothing. Nothing is deleted: sign in with the same account and everything comes back; another account sees only its own. A setup made before signing in moves into the account (unless the account has its own unsaved setup).
A sign-in lasts 60 days and extends itself with regular use. The Free plan gives 5 pages per run, 500 rows per export and 2 saved scrapers. When a new account has a trial offer, a green banner appears above Setup and a Try free chip in the top line (no card needed). × hides the banner for 7 days; the chip stays. If a run is stopped by your plan's limits and the trial would allow it, press Start free trial: the trial starts and the run continues by itself. Prices come from the server; during an offer the full price is crossed out with the discounted price beside it.
| What | Plan needed |
|---|---|
| 50 pages / 5,000 rows / 10 scrapers, product pages | Starter |
| Unlimited pages and rows, 50 scrapers, click mode, the site's API, embedded JSON | Pro |
| Image download, proxies, unlimited scrapers | Max |
- A control that isn't in your plan has a small badge saying which plan it needs. Trying to run it shows a message in Setup and an Upgrade button; the run doesn't start, and nothing is opened or clicked.
- If an export has more rows than your limit, the first N rows are exported and that is noted; Results keeps them all, so after upgrading you don't need to run again.
- Upgrade → Starter / Pro / Max, Monthly or Yearly → checkout opens in a new tab. After payment the card shows "You're on …" by itself. For invoices, your card or cancelling, use Manage billing.
- One account, one computer. Signing in elsewhere stops runs here, with the message "Your Rowvana account is now in use on another device"; Sign in here brings it back.
- Offline, your plan keeps working for up to 24 hours after the last check, then falls back to Free; when you're back online, your plan returns.
- To delete your account: the red card at the very bottom of ⚙. The Delete button turns on only after you type
Delete my Rowvana accountin the box. Cancel your plan in Manage billing first. That account's scrapers, results and settings in this browser are deleted too. Reset only signs you out; it doesn't delete the account.
Common problems
All popup buttons are grey and it says "This page cannot be scraped"
chrome://, the Web Store, a new tab). Open a normal website.Extract gives 0 rows
One column is empty, the others are fine
src and links need href, not text. Or some cards don't have that item (not every product has a discount), which is normal.Pagination is on, but only the first page comes
Infinite scroll is on, but the whole list doesn't scroll
I pressed ▶ on an old setup and got 2 rows
On Maps I got only 2 rows, then it stopped
/place/...), not the search results page. Go to the results list and run ⚡ Auto-setup again; it turns on infinite scroll by itself when needed.Every row on Maps has the same data
_detail_error instead of copied data. If you still see this, run ⚡ Auto-setup again, and if it continues, write to support@rowvana.com with the site's link.The Phone column shows the address
data-item-id and aria-label. With an older saved profile, run Auto-setup again.Odd columns in Specifications ("LIVE", "Sep 19, 12")
Non-English text looks broken in Excel
Does closing the popup stop the job?
The picker panel covers the page
Still stuck?
Good habits
- Preview first, then Extract. Check one page looks right before letting it loose on 50.
- Don't lower the delays. 700 ms per page, 1000 ms per product; below that the site may block your IP.
- Split big jobs. For 500+ products, run in rounds with Max items 100 and export each time.
- Save your setup as a scraper in the Saved tab when you're done. If the site changes, you only fix the broken field.
- Follow the site's rules. Check its Terms of Service and robots.txt, and think about how you use personal data (phone numbers, emails).
- If a login is needed, log in yourself first. Rowvana runs in your browser's session, so nothing else is needed.