GitHub Data Connector Deployment Guide
Production operating guide for the GitHub data connector covering authentication, GitHub API rate limits, and operational tuning.
Authentication & Secrets​
The GitHub connector uses the GitHub REST and GraphQL APIs with a personal access token (PAT) or GitHub App installation token.
| Parameter | Description |
|---|---|
github_token | PAT or installation token. Use ${secrets:...} to resolve from a secret store. |
Tokens must be sourced from a secret store in production. Scope the PAT to the minimum required permissions:
- Public repo data only: no token required, but see the rate-limit note below.
- Private repos:
reposcope. - Issues/PRs:
repo(private) orpublic_repo(public). - Org-level data:
read:org.
For long-running deployments, prefer GitHub App tokens (installation tokens) over user PATs — they have higher rate limits (15,000/hr vs 5,000/hr per authenticated user) and are not tied to a specific user account.
Resilience Controls​
Rate Limiting​
GitHub's REST API rate limits:
| Auth mode | Limit |
|---|---|
| Unauthenticated | 60 requests/hr per IP |
| Authenticated (PAT) | 5,000 requests/hr |
| GitHub App installation | 15,000 requests/hr |
| Enterprise Server (typical) | Configurable |
The connector respects GitHub's Retry-After and X-RateLimit-Reset headers and backs off accordingly. It stops issuing requests while the primary limit's remaining budget is at or below 10% of the limit, so other users of the same token keep a reserve, and resumes at the next reset window.
GraphQL is paced separately. GitHub's GraphQL secondary rate limit is 2,000 points per minute, and a non-mutation GraphQL request costs 1 point regardless of how many nodes or nested connections it asks for. The connector paces itself to 90% of that — a sustained 1,800 requests per minute — and every GraphQL-backed table on one authentication context shares a single limiter, so adding datasets divides that budget rather than multiplying it. Because every query costs the same 1 point, the shared limiter is first-come-first-served across tables: one large scan cannot starve the others. A secondary-limit 403 is retried on the same page using the response's retry-after.
GraphQL CPU time is not estimated locally: HTTP duration is not GitHub's CPU accounting, and budgeting against it would serialize scans GitHub would still accept.
Pagination​
Page width is chosen per table, not fixed at GitHub's 100-item maximum. GitHub enforces a per-request compute budget (GraphQL API resource limits announcement) that is separate from — and reached long before — the 500,000-node ceiling, and a page wide enough to exceed it is rejected outright with Resource limits for this query exceeded, every node in the page returned as null. Tables whose rows expand into many nested connections therefore request narrower pages: pulls is requested 25 at a time in both comment modes, while issues and milestones still use 100.
Nested connections are paged too. reviews, review_threads and release_assets exist only underneath a parent in GitHub's GraphQL schema, and GitHub caps a nested connection page at 100 nodes and will not page it from inside the parent query. The connector fetches the remainder with follow-up requests addressed at the parent, so a pull request with more than 100 reviews is read completely rather than truncated — a truncated nested connection would drop whole rows, and a short COUNT(*) gives no sign it is short. When a nested connection cannot be continued, the scan fails and names the parent rather than returning a partial set.
Datasets backed by high-volume endpoints (e.g., repos.commits on a monorepo) may require many hours to initially hydrate. Use incremental acceleration with a since filter where possible.
pulls scan ceilings differ by query modeIn github_query_mode: auto, the connector fails with Maximum pagination iterations (1000) exceeded after 1,000 pagination iterations (1,001 total page fetches counting the initial page), rather than silently truncating. At 25 rows per page, that bounds a pulls dataset at 25 x 1001 = 25,025 pull requests per refresh. In github_query_mode: search, GitHub Search has its own 1,000-result limit, which is the effective cap for a single query.
The page was wider in earlier releases, putting the arithmetic ceiling at 100,100 — but on a repository large enough for that to bind, the wider page was rejected by the compute budget and returned no rows at all. The number of pull requests actually reachable went up, not down.
Retry Behavior​
Transient 5xx responses are retried with exponential backoff up to a bounded retry count. Permanent errors (401 Unauthorized, 404 Not Found, 422 Validation Failed) surface immediately.
Capacity & Sizing​
- Throughput: Bounded by the rate limit, not network or CPU. Plan dataset refresh intervals to stay within the hourly budget.
- Latency: Expect ~100-500ms per paginated request against
github.com; lower for GitHub Enterprise Server on the same network. - Initial bootstrap: For high-volume datasets (e.g., all commits in a busy monorepo), the first materialization may exhaust the hourly budget across several runs. Plan staged ingestion if needed.
Metrics​
The GitHub connector does not register connector-specific dataset-level instruments in the current release. Monitor via:
- Spice query execution metrics (
query_duration_ms,query_returned_rows,query_failures) fromruntime.metrics. - GitHub's own rate-limit UI at
/settings/tokensfor token-level quota tracking.
See Component Metrics for general configuration.
Task History​
GitHub API calls participate in task history through the HTTP client's span. Each page fetch is a child of the enclosing sql_query or acceleration_refresh task.
Known Limitations​
- Read-only: The connector is read-only; writes (issue creation, PR comments) are not supported.
- GraphQL-only endpoints: Some GitHub data (e.g., discussions, project v2) requires GraphQL; check the connector's documented supported endpoints.
- GitHub Enterprise Cloud with IP allowlisting: The Spice runtime's outbound IP must be allow-listed.
- Secondary rate limits: GitHub enforces abuse-detection "secondary" rate limits on concentrated bursts, independent of the hourly primary limit. If hit, the connector backs off.
Troubleshooting​
| Symptom | Likely cause | Resolution |
|---|---|---|
401 Bad credentials | PAT expired / revoked / wrong value. | Rotate the PAT; update the secret store. |
403 rate limit exceeded | Primary hourly rate limit hit. | Increase refresh interval; switch to GitHub App auth for higher quota; use incremental refresh with since. |
403 Secondary rate limit | Burst of concurrent requests tripped abuse detection. | Reduce concurrent refresh; connector will back off automatically. |
404 Not Found on a private repo | Token lacks repo scope. | Regenerate PAT with repo scope. |
| Very slow initial hydration | Large dataset + strict rate limit. | Run first refresh off-peak; use since/updated_since for incremental refreshes. |
