Skip to content
PyronetsPyronets
FAQ

Frequently asked questions

Everything you need to know about how Pyronets works, from technical capabilities and data delivery to pricing and legal considerations.

Where we draw the line

The two answers people are usually checking for.

  • Bot-managed sites
  • Signed / obfuscated APIs
  • Mobile-only APIs
  • Geo-varying content
  • Anything behind someone else's login
  • Personal data without a lawful basis
What we take on

Do I need technical knowledge?

No, Pyronets is fully managed. Just tell us what data you need.

How fast can you deliver?

Sample data within 3–5 days. Full pipeline live in 1–2 weeks.

What formats do you support?

CSV, JSON, Excel, XML, API, SFTP, S3, Google Sheets, database.

Can you scrape behind login?

Yes, we handle login-based sites, CAPTCHA, and JS rendering.

How is pricing determined?

Based on volume, sources, complexity, and frequency. Free discovery call.

Is there a contract?

Flexible terms, monthly, project-based, or long-term retainer.

Questions

Getting Started

We can scrape the vast majority of publicly accessible websites, including those that are JavaScript-rendered, paginated, or require session management. Websites with login-only content can be handled if the client provides access credentials under a lawful basis. The only sites we cannot scrape are those explicitly protected by law or contractual restrictions that apply to the client.

Not at all. Our clients range from data engineers to business analysts to product managers. We handle all the technical implementation. You only need to know what data you want, from which websites, and how often. Everything else (extraction, normalization, QA, delivery, and maintenance) is handled by our team.

For most projects, we deliver a sample dataset within 3–7 business days of the discovery call. Production pipelines typically go live within 2–3 weeks depending on the number of sources and the technical complexity of the target sites. Rush timelines can sometimes be accommodated. Let us know your deadline on the discovery call.

Yes. A sample dataset is always part of the engagement process and is provided free of charge, with no obligation. The sample typically contains 500–2,000 records from your target sources and is used to validate data quality and schema before we agree on a production scope.

Questions

Data & Delivery

We support CSV, Excel (.xlsx), JSON, XML, Parquet, and custom schemas. Data can be delivered via SFTP, Amazon S3, Google Cloud Storage, Azure Blob, REST API endpoint, direct database insert, or Google Sheets. Most format and delivery combinations are available at no additional cost.

Delivery frequency ranges from real-time streaming to hourly, daily, weekly, or custom schedules. We'll recommend the appropriate frequency based on how often the source sites update and your operational requirements. More frequent delivery is always technically possible, the key question is whether the source data actually changes that often.

Yes. We support direct insert to PostgreSQL, MySQL, BigQuery, Snowflake, Redshift, and other SQL/cloud warehouse systems. Schema mapping, type coercion, and incremental upserts are all handled on our side. There is typically a one-time setup cost for database integration depending on complexity.

In many cases, yes. Historical data availability depends on the target sources. Where it exists (either on the live site via date filters or via public archives) we can extract it as a one-time backfill. Historical extraction is scoped and priced separately from ongoing pipeline pricing.

Yes. Cross-source record matching (also called entity resolution or SKU matching) is one of our core capabilities. We use GTIN, MPN, EAN, and fuzzy title/attribute matching to link product records across sources. Match confidence scores are included in the output so your team can set acceptance thresholds.

Questions

Technical

Yes, provided the client has a lawful right to access the data behind the login and provides us with valid credentials or an API token. We build session-managed scrapers that handle authentication, session refresh, and two-factor authentication where applicable. Responsibility for ensuring access is authorized rests with the client.

Yes. We operate premium residential and datacenter proxy networks and use browser automation with human-emulation techniques to handle most commercial anti-bot systems. Some extremely aggressive targets may require additional lead time or cost. We assess anti-bot complexity during the scoping phase and will flag it in the proposal.

Our monitoring systems detect structural changes on target websites (including layout updates, class name changes, and new pagination patterns) and alert the engineering team automatically. In most cases, issues are diagnosed and pipelines are restored within 24 hours of detection, often before the client notices a gap in delivery.

Yes. We use headless browser automation (Chromium-based) for sites that require JavaScript rendering, dynamic content loading, or client-side routing. This applies to React, Angular, Vue, and other SPA frameworks. Performance is slightly lower than for static-rendered pages, so this is factored into our capacity planning and pricing.

Questions

Quality & Accuracy

Accuracy is enforced at three levels: field-level validation during extraction (type checking, regex patterns, expected value ranges), cross-record consistency checks (deduplication, outlier detection, cross-source comparisons), and a final human review by a QA analyst on a random sample before each delivery. Every delivery is accompanied by a quality report showing pass rates per check.

Our standard validation suite runs 48 checks: null/missing field rates, data type conformance, URL validity, numeric range validation, date format consistency, duplicate detection, cross-source reconciliation where applicable, and category/taxonomy mapping accuracy. Clients can also define custom validation rules specific to their schema.

Our delivery reports include field completeness rates for every key field in the schema. If completeness falls below agreed thresholds, we flag it and investigate before delivery rather than shipping degraded data silently. For critical fields, we'll hold a delivery and communicate rather than let incomplete data flow into your systems.

Questions

Commercial

Pricing is based on several variables: number of target websites, total data volume, crawl frequency, technical complexity of the sources, anti-bot difficulty, number of data fields, delivery format, QA requirements, and whether historical data is needed. We provide fixed-price proposals after a discovery call and a sample dataset, so you know the full cost before committing.

Yes. Change requests are handled as part of the ongoing engagement. Adding fields, new sources, or changing delivery frequency is scoped and implemented, typically within a few days for minor changes. Your PM will advise on any impact to timeline or pricing before changes are made.

Scraping publicly available data is generally lawful in most jurisdictions, and has been affirmed in multiple court decisions including hiQ Labs v. LinkedIn. However, legality depends on how the data is used, the jurisdiction, the specific website's terms of service, and data protection regulations like GDPR. Pyronets scrapes only publicly accessible data. Clients are responsible for ensuring that their collection, storage, and use of the data complies with all applicable laws and regulations in their jurisdiction.

We adapt to your preferred tools. Most clients use a combination of email and Slack for day-to-day communication, Basecamp or Notion for project documentation, and a weekly or bi-weekly check-in call for larger ongoing projects. Your PM will set up the right channels during onboarding.

Still have questions?

Send us a message and we'll get back to you within one business day.

We'll only use these details to reply about your enquiry.

Ready to see the data for yourself?

Request a free sample dataset, no commitment, no contract. Just real data from your target sources.

No retainer required to find out whether your sources are feasible.