Providers sit at the center
Provider SDKs account for 25.0% of all normalized edges, and OpenAI SDK appears in 504 final repos. The provider layer is the most common and the most connective part of the stack.
+--------------------------------------------------------------+ | OSS AI STACK MAP :: SNAPSHOT REPORT | | SOURCE: data/run-2026-07-30-repaired-v3 | | LENS: major, active, public OSS AI repos on GitHub | +--------------------------------------------------------------+
This report reads directly from data/run-2026-07-30-repaired-v3 and summarizes the current stack choices across the project’s final GitHub AI set.
Study frame: GitHub-only, public, non-fork, non-archived, active within 1 month, and at least 1,000 stars. Published stack edges come from manifests, SBOMs, bounded import fallback, repo identity, and explicitly approved README fallback when an included repo would otherwise remain unmapped.
Provider SDKs account for 25.0% of all normalized edges, and OpenAI SDK appears in 504 final repos. The provider layer is the most common and the most connective part of the stack.
299 of 994 technology-mapped repos (30.1%) use at least two tracked providers. The most common provider pairing is Anthropic SDK plus OpenAI SDK in 233 repos. That is followed by Google GenAI SDK plus OpenAI SDK in 194 repos. Major OSS projects are not clustering around a single vendor.
Training, orchestration, providers, and retrieval dominate the graph. Evaluation and observability remain comparatively thin, with only 28 guardrail/eval edges and 67 observability edges.
Inference from the aggregate counts: the modal major OSS AI repo in this snapshot is organization-owned, Python-first, anchored on a provider SDK, often layers in Hugging Face training tools, and then adds orchestration, retrieval, and a lightweight UI shell.
This layer is separate from stack normalization. Repo steward mapping currently uses curated exact repo-name and GitHub org matches only. Technology vendor mapping is curated in config and should be read as product stewardship, not proof that every adopting repo is company-backed.
1,032 final repos have manifests and 816 have SBOM dependency evidence. 12 repos map only via canonical repo identity, 0 combine direct evidence with fallback signals, and 0 remain README-only.
49 included repos (4.7%) have no normalized technology edge, so graph-like analysis describes the mapped subset, not the entire final population.
0 repos were judge-reviewed and 0 judge overrides were applied in this snapshot. The published set remains rule-driven. A judge-backed validation audit was not run for this snapshot; no false-positive estimate is available.
0 repos record a material fetch or parse failure. 1,504 legacy collection outcomes remain unknown because this repaired snapshot predates per-source outcome tracking.
Rule-only yields 1,043 final repos. Judge adjustment yields 1,043. Direct-only evidence maps 994 repos, and approved fallback lifts that to 994. No temporal baseline comparison is available for this snapshot.
These cards rank categories by normalized repo-tech edges, not by architectural importance. Each card now shows both edge share and repo prevalence across the 1,043 final repos. Very thin categories are summarized separately below.
These visuals summarize the technology-connected subset of the final population. Eigenvector highlights the core hubs, betweenness isolates bridge technologies, repo degree shows stack breadth per mapped repo, and category mixing shows which layers of the stack actually co-occur.
| Technology | Eigenvector centrality |
|---|---|
| OpenAI SDK | 0.4231 |
| Model Context Protocol | 0.3326 |
| Anthropic SDK | 0.3051 |
| Transformers | 0.3049 |
| PyTorch | 0.2691 |
| LangChain | 0.2377 |
| Hugging Face Hub | 0.2282 |
| Google GenAI SDK | 0.2253 |
| LiteLLM | 0.1904 |
| LangChain OpenAI Integration | 0.1811 |
| Tokenizers | 0.1719 |
| Accelerate | 0.1701 |
| Technology | Repo count | Betweenness | Weighted strength |
|---|---|---|---|
| OpenAI SDK | 523 | 0.7173 | 2,999 |
| Transformers | 314 | 0.1859 | 2,034 |
| Model Context Protocol | 482 | 0.0732 | 2,190 |
| LangChain | 199 | 0.0500 | 1,649 |
| Vercel AI SDK | 124 | 0.0253 | 709 |
| Hugging Face Hub | 195 | 0.0253 | 1,470 |
| Qdrant | 71 | 0.0242 | 742 |
| PyTorch | 270 | 0.0038 | 1,774 |
| Anthropic SDK | 303 | 0.0023 | 1,932 |
| LangChain OpenAI Integration | 138 | 0.0011 | 1,251 |
| PEFT | 89 | 0.0005 | 753 |
| Milvus | 36 | 0.0004 | 346 |
| Browserbase | 24 | 0.0004 | 211 |
| LangChain Google GenAI Integration | 47 | 0.0004 | 410 |
| LiteLLM | 153 | 0.0003 | 1,292 |
| Daytona | 29 | 0.0002 | 280 |
| Streamlit | 66 | 0.0002 | 591 |
| Modal | 22 | 0.0002 | 234 |
| Accelerate | 132 | 0.0002 | 1,092 |
| LangGraph | 105 | 0.0001 | 1,047 |
| Tokenizers | 141 | 0.0001 | 1,122 |
| vLLM | 52 | 0.0000 | 458 |
| Chroma | 74 | 0.0000 | 786 |
| Langfuse | 43 | 0.0000 | 347 |
| E2B | 45 | 0.0000 | 343 |
| Weaviate | 39 | 0.0000 | 380 |
| pgvector | 46 | 0.0000 | 358 |
| code2prompt | 1 | 0.0000 | 0 |
| Gradio | 74 | 0.0000 | 620 |
| LlamaIndex | 42 | 0.0000 | 430 |
| Browserless | 1 | 0.0000 | 0 |
| Arize Phoenix | 4 | 0.0000 | 48 |
| Ray Serve | 46 | 0.0000 | 355 |
| Google ADK | 26 | 0.0000 | 286 |
| Promptfoo | 3 | 0.0000 | 21 |
| OpenAI Agents | 41 | 0.0000 | 452 |
| DeepSpeed | 32 | 0.0000 | 237 |
| Semantic Kernel | 8 | 0.0000 | 102 |
| Helicone | 1 | 0.0000 | 12 |
| Guardrails | 4 | 0.0000 | 38 |
| Chainlit | 10 | 0.0000 | 112 |
| Browser Use | 11 | 0.0000 | 89 |
| DSPy | 11 | 0.0000 | 130 |
| DeepEval | 7 | 0.0000 | 73 |
| LanceDB | 45 | 0.0000 | 424 |
| CrewAI | 34 | 0.0000 | 451 |
| Weave | 2 | 0.0000 | 23 |
| Grafbase | 1 | 0.0000 | 0 |
| SGLang | 12 | 0.0000 | 125 |
| Logfire | 16 | 0.0000 | 134 |
| TRL | 19 | 0.0000 | 176 |
| Vercel Sandbox | 7 | 0.0000 | 45 |
| LangChain Anthropic Integration | 64 | 0.0000 | 612 |
| Hyperbrowser | 3 | 0.0000 | 25 |
| DingoDB | 1 | 0.0000 | 0 |
| PrimeIntellect Verifiers | 1 | 0.0000 | 14 |
| llama.cpp | 22 | 0.0000 | 240 |
| TGI | 1 | 0.0000 | 7 |
| AutoGen | 6 | 0.0000 | 95 |
| Haystack | 1 | 0.0000 | 1 |
| smolagents | 9 | 0.0000 | 115 |
| Cloudflare Agents | 5 | 0.0000 | 25 |
| BentoML | 2 | 0.0000 | 10 |
| Ollama | 93 | 0.0000 | 700 |
| NeMo Guardrails | 2 | 0.0000 | 26 |
| Google GenAI SDK | 198 | 0.0000 | 1,408 |
| Cloudflare Containers | 5 | 0.0000 | 24 |
| Evidently | 1 | 0.0000 | 0 |
| Mastra | 18 | 0.0000 | 229 |
| Runloop | 3 | 0.0000 | 24 |
| Instructor | 25 | 0.0000 | 302 |
| Ragas | 12 | 0.0000 | 132 |
| PydanticAI | 15 | 0.0000 | 212 |
| Notte | 1 | 0.0000 | 6 |
| Steel Browser | 2 | 0.0000 | 14 |
| Databend | 3 | 0.0000 | 10 |
| Tracked technologies per repo | Repo count |
|---|---|
| 1 | 248 |
| 2 | 148 |
| 3 | 112 |
| 4 | 95 |
| 5 | 81 |
| 6-7 | 104 |
| 8-10 | 105 |
| 11-15 | 73 |
| 16+ | 28 |
| Category A | Category B | Weighted co-occurrence |
|---|---|---|
| Model frameworks and HF stack | Model frameworks and HF stack | 1,723 |
| Model frameworks and HF stack | Providers and access | 1,619 |
| Model frameworks and HF stack | Orchestration and agents | 1,119 |
| Model frameworks and HF stack | Protocols and developer SDKs | 486 |
| Model frameworks and HF stack | Retrieval and vector storage | 572 |
| Model frameworks and HF stack | Serving and local runtimes | 711 |
| Model frameworks and HF stack | UI and app frameworks | 422 |
| Model frameworks and HF stack | Sandbox and isolated execution | 141 |
| Providers and access | Providers and access | 793 |
| Providers and access | Orchestration and agents | 1,640 |
| Providers and access | Protocols and developer SDKs | 930 |
| Providers and access | Retrieval and vector storage | 676 |
| Providers and access | Serving and local runtimes | 387 |
| Providers and access | UI and app frameworks | 256 |
| Providers and access | Sandbox and isolated execution | 232 |
| Orchestration and agents | Orchestration and agents | 1,342 |
| Orchestration and agents | Protocols and developer SDKs | 672 |
| Orchestration and agents | Retrieval and vector storage | 615 |
| Orchestration and agents | Serving and local runtimes | 265 |
| Orchestration and agents | UI and app frameworks | 272 |
| Orchestration and agents | Sandbox and isolated execution | 182 |
| Protocols and developer SDKs | Protocols and developer SDKs | 77 |
| Protocols and developer SDKs | Retrieval and vector storage | 217 |
| Protocols and developer SDKs | Serving and local runtimes | 129 |
| Protocols and developer SDKs | UI and app frameworks | 74 |
| Protocols and developer SDKs | Sandbox and isolated execution | 112 |
| Retrieval and vector storage | Retrieval and vector storage | 255 |
| Retrieval and vector storage | Serving and local runtimes | 139 |
| Retrieval and vector storage | UI and app frameworks | 112 |
| Retrieval and vector storage | Sandbox and isolated execution | 88 |
| Serving and local runtimes | Serving and local runtimes | 60 |
| Serving and local runtimes | UI and app frameworks | 80 |
| Serving and local runtimes | Sandbox and isolated execution | 29 |
| UI and app frameworks | UI and app frameworks | 20 |
| UI and app frameworks | Sandbox and isolated execution | 25 |
| Sandbox and isolated execution | Sandbox and isolated execution | 45 |
This section compares the current 2026-07-30 snapshot against the earlier published 2026-03-25 snapshot (run-2026-03-31-publication-v8). Shares are of each run's own final repo set; the scorecard marks benchmark recall and coverage metrics as improved or regressed. The LLM-judge hardening layer ran in one run but not the other; the 2026-07-30 pass is rule-only, so its serious/AI-relevance filtering relies on rule scores alone — the “Judge-reviewed repos” scorecard row reflects this asymmetry, not a real coverage change.