The invisible layer

Research data workflows run on open source:

  • Data collection, cleaning, analysis, visualization
  • Reproducibility tooling: containers, workflows, version control
  • Repositories and scholarly infrastructure

92% of researchers use software; 69% say their research would not be practical without it.

The people keeping it running are increasingly called Research Software Engineers (RSEs), a growing professional community with no standard institutional home.

If every organization had to replace open source with proprietary equivalents, the cost would exceed $8.8 trillion.

Hettrick et al. (SSI, 2014/2022) · Hoffmann, Nagle & Zhou (2024) doi:10.2139/ssrn.4693148

JupyterJupyter
PythonPython
RR
GitGit
Apache SparkSpark
PandasPandas
NumPyNumPy
DockerDocker
RayRay
ggplot2ggplot2
dplyrdplyr
SnakemakeSnakemake
NextflowNextflow

The UC research-to-infrastructure pipeline

Data professionals built open data infrastructure: FAIR, DMPs, curation, provenance.

FAIR4RS is the same inflection point for software — and UC has been showing the way for 50 years.

BSD Unix · Ceph · Apache Spark · Jupyter · RISC-V · Ray

The pipeline from university research to global infrastructure is not rare or accidental.

Barker et al. (2022) doi:10.1038/s41597-022-01710-x

UC Berkeley
JupyterJupyter
Apache SparkApache Spark
RISC-VRISC-V
RayRay
BSD UnixBSD Unix
UC Santa Cruz
CephCeph

UCSC Genome Browser

UC San Diego
Cytoscape EEGLAB
UCSD Pascal p-SystemUCSD Pascal
UCLA
ProcessingProcessing
Named Data NetworkingNamed Data Networking

tz database

UC San Francisco
ChimeraX
OpenMMOpenMM
UC Davis
sourmashsourmash

ERPLAB

What is an academic OSPO?

An Open Source Program Office supports, governs, and promotes open source within an institution.

  • Originated in industry: ~77% of large tech companies have one
  • First academic OSPOs: Johns Hopkins, UC Santa Cruz, and RIT (2021)
  • 34 academic OSPOs in CURIOSS worldwide and growing
  • Core functions: policy, licensing guidance, best practices, education, community

🟠 UC OSPO Network  ·  🔵 CURIOSS member

World map showing 34 CURIOSS member institutions. Six UC campuses (orange) in California: Santa Cruz, Berkeley, Davis, Los Angeles, Santa Barbara, and San Diego. Twenty-eight additional institutions (blue) across the US, UK, Switzerland, Spain, Ireland, Luxembourg, and France.

The UC OSPO Network

A multi-campus collaboration treating open source as shared infrastructure.

  • 6 campuses: Santa Cruz, Berkeley, Davis, Los Angeles, Santa Barbara, San Diego
  • Launched April 2024, funded by the Alfred P. Sloan Foundation
  • $1.85M in external funding to date
  • Lead: UC Santa Cruz, the first OSPO in a large state university system
  • Serves 280,000+ students and 25,000+ faculty

Acts as a neutral convener: no single campus, company, or grant cycle can pull the work away.

Ruff (2026) UC Open Summit · youtu.be/eBriL3CDNeo

Map of California showing six UC OSPO Network campuses: Santa Cruz, Berkeley, Davis, Los Angeles, Santa Barbara, and San Diego.

Discovery · Sustainability · Education

🔭

Discovery

Mapping who does what across the UC system

  • 236,000+ repos scanned; ~82,000 institutionally affiliated
  • 294-respondent system-wide survey
  • UC Open Repository Browser (UC ORB)

🌱

Sustainability

Keeping projects and communities healthy

  • 58% of UC contributors are also maintainers
  • Cross-campus licensing working group (UCOP + tech transfer)
  • Project health assessment (OSSPREY)

📖

Education

Coordinated training across campuses

  • Gap analysis → curriculum inventory → learning pathways
  • Published at ucospo.net/education
  • Coordinated with Carpentries infrastructure

Gomez et al. (2025) arxiv:2506.18359 · Scarlett et al. (2025) doi:10.31235/osf.io/p8bx6_v1

Lessons from two years

What the network model enables that single campuses cannot:

  1. Shared staffing: community manager, licensing specialist, technical roles
  2. Policy leverage: the network has standing with UCOP; individual campuses do not
  3. Data at scale: the GitHub and survey analyses require system-wide scope to mean anything
  4. Equity: smaller campuses get services they could not build alone
  5. A replicable model: $1.85M produced something other state systems can follow

Each campus contributes unique expertise; the network routes it to where it’s needed.

Where this meets your work

For data professionals, the OSPO sits at the intersection of:

  • Research software engineering ↔︎ data management
  • Licensing compliance ↔︎ open data policy
  • Reproducibility tooling ↔︎ FAIR / FAIR4RS
  • Training infrastructure ↔︎ data literacy programs
  • RSE community ↔︎ library research support

For your institution:

  • Where does open source software governance currently live?
  • What would it take to treat it as infrastructure rather than a side project?

You’re already doing this work

If you do technical data services, researchers are already bringing you these questions:

  • “What license should I use for this code?”
  • “Is this library still maintained? Should I depend on it?”
  • “How do I make this reproducible and citable?”
  • “My funder wants a software management plan.”
  • “How do I accept contributions from collaborators?”

That is OSPO work. The frame makes it legible to your institution.

The network has already built the resources (Laura Langdon, UC OSPO):

Raise the floor, not just solve the ticket

Fixing their immediate problem gets them through the week.

Moving them from “it runs on my machine” to maintainable, licensed, citable software gets their whole lab there, and their students after them.

You don’t need a formal OSPO to do this. You need the frame.

OSPO heuristics already in your toolkit: health signals · bus factor · governance files · FAIR4RS · software DMPs · contributor conventions · community governance

Learn more / connect

UC OSPO Network: ucospo.net

Global network: curioss.org, 34 academic OSPOs worldwide and growing

Tim Dennis Data Science Center, UCLA Library ORCIDorcid.org/0000-0001-6632-3812 tdennis@library.ucla.edu

QR code linking to tinyurl.com/iassist2026 — slides for this presentation

tinyurl.com/iassist2026

References

Barker, M., Chue Hong, N. P., Katz, D. S., Lamprecht, A.-L., Martinez-Ortiz, C., Psomopoulos, F., Harrow, J., Castro, L. J., Gruenpeter, M., Martinez, P. A., & Honeyman, T. (2022). Introducing the FAIR principles for research software. Scientific Data, 9, 622. https://doi.org/10.1038/s41597-022-01710-x
Brown, C. T., & Irber, L. (2016). Sourmash: A library for MinHash sketching of DNA. Journal of Open Source Software, 1(5), 27. https://doi.org/10.21105/joss.00027
Cerf, V. G., & Kahn, R. E. (1974). A protocol for packet network intercommunication. IEEE Transactions on Communications, 22(5), 637–648. https://doi.org/10.1109/TCOM.1974.1092259
Delorme, A., & Makeig, S. (2004). EEGLAB: An open source toolbox for analysis of single-trial EEG dynamics including independent component analysis. Journal of Neuroscience Methods, 134(1), 9–21. https://doi.org/10.1016/j.jneumeth.2003.10.009
Eastman, P., Swails, J., Chodera, J. D., McGibbon, R. T., Zhao, Y., Beauchamp, K. A., Wang, L.-P., Simmonett, A. C., Harrigan, M. P., Stern, C. D., Wiewiora, R. P., Brooks, B. R., & Pande, V. S. (2017). OpenMM 7: Rapid development of high performance algorithms for molecular dynamics. PLOS Computational Biology, 13(7), e1005659. https://doi.org/10.1371/journal.pcbi.1005659
Gomez, J., Lovell, E., Lieggi, S., Cardenas, A. A., & Davis, J. (2025). Recipe for discovery: A pipeline for institutional open source activity. https://arxiv.org/abs/2506.18359
Hettrick, S., Antonioletti, M., Carr, L., Chue Hong, N., Crouch, S., De Roure, D., Emsley, I., Goble, C., Hay, A., Inupakutika, D., et al. (2014). UK Research Software Survey 2014. Software Sustainability Institute. https://doi.org/10.5281/zenodo.14809
Hoffmann, M., Nagle, F., & Zhou, Y. (2024). The value of open source software (Working Paper Nos. 24-038). Harvard Business School. https://doi.org/10.2139/ssrn.4693148
Kent, W. J., Sugnet, C. W., Furey, T. S., Roskin, K. M., Pringle, T. H., Zahler, A. M., & Haussler, D. (2002). The human genome browser at UCSC. Genome Research, 12(6), 996–1006. https://doi.org/10.1101/gr.229102
Kluyver, T., Ragan-Kelley, B., Pérez, F., Granger, B., Bussonnier, M., Frederic, J., Kelley, K., Hamrick, J., Grout, J., Corlay, S., Ivanov, P., Avila, D., Abdalla, S., & Willing, C. (2016). Jupyter Notebooks — a publishing format for reproducible computational workflows. In F. Loizides & B. Schmidt (Eds.), Positioning and power in academic publishing: Players, agents and agendas (pp. 87–90). IOS Press. https://doi.org/10.3233/978-1-61499-649-1-87
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Elibol, M., Yang, Z., Paul, W., Jordan, M. I., & Stoica, I. (2018). Ray: A distributed framework for emerging AI applications. 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 561–577.
Olson, A. D., & Eggert, P. (1986). tz database (IANA time zone database). IANA. https://www.iana.org/time-zones
Ruff, N. (2026). The role of foundations in advancing open collaboration and innovation. UC Open Summit. https://youtu.be/eBriL3CDNeo
Scarlett, V., Curty, R. G., Gomez, J., Langdon, L., Janée, G., & Budden, A. E. (2025). A system-wide snapshot: A multi-campus survey of open source contributors at the University of California. SocArXiv. https://doi.org/10.31235/osf.io/p8bx6_v1
Shannon, P., Markiel, A., Ozier, O., Baliga, N. S., Wang, J. T., Ramage, D., Amin, N., Schwikowski, B., & Ideker, T. (2003). Cytoscape: A software environment for integrated models of biomolecular interaction networks. Genome Research, 13(11), 2498–2504. https://doi.org/10.1101/gr.1239303
Stonebraker, M., & Rowe, L. A. (1987). The design of POSTGRES (UCB/ERL M86/85). University of California, Berkeley.
Stonebraker, M., Wong, E., Kreps, P., & Held, G. (1976). The design and implementation of INGRES. ACM Transactions on Database Systems, 1(3), 189–222. https://doi.org/10.1145/320473.320476
Tidelift. (2024). The 2024 Tidelift maintainer impact report. Tidelift. https://dev.to/tidelift/the-open-source-maintainer-community-is-getting-grayer-1gc2
Waterman, A., & Asanović, K. (2019). The RISC-V instruction set manual, volume I: Unprivileged ISA. RISC-V Foundation. https://github.com/riscv/riscv-isa-manual
Weil, S. A., Brandt, S. A., Miller, E. L., Long, D. D. E., & Maltzahn, C. (2006). Ceph: A scalable, high-performance distributed file system. 7th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 307–320.
Zaharia, M., Chowdhury, M., Das, T., Dave, A., Ma, J., McCauly, M., Franklin, M. J., Shenker, S., & Stoica, I. (2012). Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI), 15–28.
Zhang, L., Afanasyev, A., Burke, J., Jacobson, V., claffy, kc, Crowley, P., Papadopoulos, C., Wang, L., & Zhang, B. (2014). Named data networking. ACM SIGCOMM Computer Communication Review, 44(3), 66–73. https://doi.org/10.1145/2656877.2656887