---
title: "Export data in bulk with Python"
description: "Run a query server-side as an asynchronous export job and download the results as files: create, poll, download, straight into pandas."
slug: "/docs/guides/bulk-export"
last_updated: "2026-08-31"
status: "published"
industry: []
location_types: []
tags: ["platform", "api", "python", "export"]
is_case_study: false
locked: false
related_metrics: []
related_datasets: []
example_count: 0
example_tools: []
---


# Export data in bulk with Python

Run a query server-side as an asynchronous export job and download the results as files: create, poll, download, straight into pandas.

## Metadata

- Tags: platform, api, python, export

Some questions end in a table too large to page through: full history for a brand, a whole metro's series, the training set for a model. The export API is built for exactly that: it runs the same query server-side as an asynchronous, high-throughput job and returns downloadable files, free of the 90-second per-request timeout that interactive queries run under.

The request shape (`dataReferenceName`, `fields`, `filters`, `filterLogic`) is identical to [`:runQuery`](/docs/guides/first-query), so anything built there ports directly. Exports return up to 100 million records by default (`limit` caps it lower; result sets beyond the cap fail with `RESOURCE_EXHAUSTED` rather than truncating silently).

---

## 1. Create the export

One POST starts the job. This example pulls a brand's June-2025 month in one city. Swap in your own filters, or none at all for a full dataset pull:

```python
import os, requests

BASE = "https://api.pine59.com"
auth = {"Authorization": f"Bearer {os.environ['PINE59_API_KEY']}"}

job = requests.post(f"{BASE}/v3/exports", headers=auth, json={
    "dataReferenceName": "visitation_usa_places.foot_traffic_month",
    "fields": [{"name": "location_id"}, {"name": "street_address"}, {"name": "visits_sum"}],
    "filters": [
        {"fieldName": "brands", "operator": "==", "value": "H-E-B"},
        {"fieldName": "city", "operator": "==", "value": "San Antonio"},
        {"fieldName": "observation_start_date", "operator": "==", "value": "2025-06-01"},
    ],
    "format": "CSV",
}).json()

print(job["id"], job["status"])   # export_..., RUNNING
```

`format` takes `CSV`, `JSON`, `AVRO`, or `PARQUET`. Parquet is the right choice for anything headed into pandas or a warehouse at volume. `fileConfig.filenamePrefix` names the output files when you organize many exports.

## 2. Poll until it completes

Exports move through `RUNNING` (and sometimes `COPYING`) to `COMPLETED` or `FAILED`:

```python
import time

while job["status"] in ("RUNNING", "COPYING"):
    time.sleep(5)
    job = requests.get(f"{BASE}/v3/exports/{job['id']}", headers=auth).json()

print(job["status"])
```

The small run above completed in under twenty seconds; full-dataset exports take longer, so poll every few seconds rather than continuously.

## 3. Download the files

A completed export lists its files under `fileExportInfo.files`, each with a `url`, `name`, `sizeBytes`, and format metadata. The URLs are time-limited (`fileExportInfo.expireTime` says until when), so download promptly and re-create the export if the links lapse:

```python
import pandas as pd
from io import BytesIO

frames = []
for f in job["fileExportInfo"]["files"]:
    content = requests.get(f["url"]).content
    frames.append(pd.read_csv(BytesIO(content)))

df = pd.concat(frames, ignore_index=True)
print(len(df), "rows")
```

For June 2025, the example above returns 45 rows, one per H-E-B store in San Antonio, ready for the same notebook work the [examples](/docs/guides) build on. Large exports split across multiple files; the loop handles either case.

## 4. Manage export jobs

The remaining endpoints round out the lifecycle: `GET /v3/exports` lists your exports, and `DELETE /v3/exports/{id}` removes one. For exports that should land directly in your own BigQuery project instead of files, the create request takes a `bigqueryConfig`. The [API reference](/docs/api) documents the destination options.

## When to export, when to query

`:runQuery` is built for interactive work: filtered questions, aggregations, anything a person or agent iterates on. As data volumes grow, the export API becomes the more effective path: the same query, run server-side at full throughput, delivered as files sized for notebooks and warehouses. The two compose naturally: explore and refine interactively, then export the refined query at scale.
