Enterprise

With $21M in funding, Code Ocean aims to help researchers replicate data-heavy science

Comment

Illustration showing scientists using the Code Ocean product as if it were a real physical thing.
Image Credits: Code Ocean

Every branch of science is increasingly reliant on big data sets and analysis, which means a growing confusion of formats and platforms — more than inconvenient, this can hinder the process of peer review and replication of research. Code Ocean hopes to make it easier for scientists to collaborate by making a flexible, shareable format and platform for any and all data sets and methods, and it has raised a total of $21 million to build it out.

Certainly there’s an air of “Too many options? Try this one!” to this (and here’s the requisite relevant XKCD). But Code Ocean isn’t creating a competitor to successful tools like Jupyter or GitLab or Docker — it’s more of a small-scale container platform that lets you wrap up all the necessary components of your data and analysis in an easily shared format, whatever platform they live on natively.

The trouble appears when you need to share what you’re doing with another researcher, whether they’re on the bench next to you or at a university across the country. It’s important for replication purposes that data analysis — just like any other scientific technique — be done exactly the same way. But there’s no guarantee that your colleague will use the same structures, formats, notation, labels and so on.

That doesn’t mean it’s impossible to share your work, but it does add a lot of extra steps as would-be replicators or iterators check and double check that all the methods are the same, that the same versions of the same tools are being used in the same order, with the same settings, and so on. A tiny inconsistency can have major repercussions down the road.

Turns out this problem is similar in a way to how many cloud services are spun up. Software deployments can be as finicky as scientific experiments, and one solution to this is containers, which like tiny virtual machines include everything needed to accomplish a computing task, in a portable format compatible with many different setups. The idea is a natural one to transfer to the research world, where you can tie up all in one tidy package the data, the software used and the specific techniques and processes used to reach a given result. That, at least, is the pitch Code Ocean offers for its platform and “Compute Capsules.”

Diagram showing how a "compute capsule" includes code, environment, and data.
Image Credits: Code Ocean

Say you’re a microbiologist looking at the effectiveness of a promising compound on certain muscle cells. You’re working in R, writing in RStudio on an Ubuntu machine, and your data are such and such collected during an in vitro observation. While you would naturally declare all this when you publish, there’s no guarantee anyone has an Ubuntu laptop with a working RStudio setup around, so even if you provide all the code, it might be for nothing.

If, however, you put it on Code Ocean, like this, it makes all the relevant code available, and capable of being inspected and run unmodified with a click, or being fiddled with if a colleague is wondering about a certain piece. It works through a single link and web app, cross platform, and can even be embedded on a webpage like a document or video. (I’m going to try to do that below, but our backend is a little finicky. The capsule itself is here.)

More than that, though, the Compute Capsule can be repurposed by others with new data and modifications. Maybe the technique you put online is a general purpose RNA sequence analysis tool that works as long as you feed it properly formatted data, and that’s something others would have had to code from scratch in order to take advantage of some platforms.

Well, they can just clone your capsule, run it with their own data and get their own results in addition to verifying your own. This can be done via the Code Ocean website or just by downloading a zip file of the whole thing and getting it running on their own computer, if they happen to have a compatible setup. A few more example capsules can be found here.

Screenshot of the Code Ocean workbench environment.
Image Credits: Code Ocean

This sort of cross-pollination of research techniques is as old as science, but modern data-heavy experimentation often ends up siloed because it can’t easily be shared and verified even though the code is technically available. That means other researchers move on, build their own thing and further reinforce the silo system.

Right now there are about 2,000 public compute capsules on Code Ocean, most of which are associated with a published paper. Most have also been used by others, either to replicate or try something new, and some, like ultra-specific open source code libraries, have been used by thousands.

Naturally there are security concerns when working with proprietary or medically sensitive data, and the enterprise product allows the whole system to run on a private cloud platform. That way it would be more of an internal tool, and at major research institutions that in itself could be quite useful.

Data is the world’s most valuable (and vulnerable) resource

Code Ocean hopes that by being as inclusive as possible in terms of codebases, platforms, compute services and so on will make for a more collaborative environment at the cutting edge.

Clearly that ambition is shared by others, as the the company has raised $21 million so far, $6 million of which was in previously undisclosed investments and $15 million in an A round announced today. The A round was led by Battery Ventures, with Digitalis Ventures, EBSCO and Vaal Partners participating as well as numerous others.

The money will allow the company to further develop, scale and promote its platform. With luck they’ll soon find themselves among the rarefied air often breathed by this sort of savvy SaaS — necessary, deeply integrated and profitable.

 

More TechCrunch

Featured Article

Industries may be ready for humanoid robots, but are the robots ready for them?

How large a role humanoids will play in that ecosystem is, perhaps, the biggest question on everyone’s mind at the moment.

30 mins ago
Industries may be ready for humanoid robots, but are the robots ready for them?

Featured Article

VCs are selling shares of hot AI companies like Anthropic and xAI to small investors in a wild SPV market

VCs are clamoring to invest in hot AI companies, willing to pay exorbitant share prices for coveted spots on their cap tables. Even so, most aren’t able to get into such deals at all. Yet, small, unknown investors, including family offices and high-net-worth individuals, have found their own way to get shares of the hottest…

2 hours ago
VCs are selling shares of hot AI companies like Anthropic and xAI to small investors in a wild SPV market

The fashion industry has a huge problem: Despite many returned items being unworn or undamaged, a lot, if not the majority, end up in the trash. An estimated 9.5 billion…

Deal Dive: How (Re)vive grew 10x last year by helping retailers recycle and sell returned items

Tumblr officially shut down “Tips,” an opt-in feature where creators could receive one-time payments from their followers.  As of today, the tipping icon has automatically disappeared from all posts and…

You can no longer use Tumblr’s tipping feature 

Generative AI improvements are increasingly being made through data curation and collection — not architectural — improvements. Big Tech has an advantage.

AI training data has a price tag that only Big Tech can afford

Keeping up with an industry as fast-moving as AI is a tall order. So until an AI can do it for you, here’s a handy roundup of recent stories in the world…

This Week in AI: Can we (and could we ever) trust OpenAI?

Jasper Health, a cancer care platform startup, laid off a substantial part of its workforce, TechCrunch has learned.

General Catalyst-backed Jasper Health lays off staff

Featured Article

Live Nation confirms Ticketmaster was hacked, says personal information stolen in data breach

Live Nation says its Ticketmaster subsidiary was hacked. A hacker claims to be selling 560 million customer records.

20 hours ago
Live Nation confirms Ticketmaster was hacked, says personal information stolen in data breach

Featured Article

Inside EV startup Fisker’s collapse: how the company crumbled under its founders’ whims

An autonomous pod. A solid-state battery-powered sports car. An electric pickup truck. A convertible grand tourer EV with up to 600 miles of range. A “fully connected mobility device” for young urban innovators to be built by Foxconn and priced under $30,000. The next Popemobile. Over the past eight years, famed vehicle designer Henrik Fisker…

21 hours ago
Inside EV startup Fisker’s collapse: how the company crumbled under its founders’ whims

Late Friday afternoon, a time window companies usually reserve for unflattering disclosures, AI startup Hugging Face said that its security team earlier this week detected “unauthorized access” to Spaces, Hugging…

Hugging Face says it detected ‘unauthorized access’ to its AI model hosting platform

Featured Article

Hacked, leaked, exposed: Why you should never use stalkerware apps

Using stalkerware is creepy, unethical, potentially illegal, and puts your data and that of your loved ones in danger.

21 hours ago
Hacked, leaked, exposed: Why you should never use stalkerware apps

The design brief was simple: each grind and dry cycle had to be completed before breakfast. Here’s how Mill made it happen.

Mill’s redesigned food waste bin really is faster and quieter than before

Google is embarrassed about its AI Overviews, too. After a deluge of dunks and memes over the past week, which cracked on the poor quality and outright misinformation that arose…

Google admits its AI Overviews need work, but we’re all helping it beta test

Welcome to Startups Weekly — Haje‘s weekly recap of everything you can’t miss from the world of startups. Sign up here to get it in your inbox every Friday. In…

Startups Weekly: Musk raises $6B for AI and the fintech dominoes are falling

The product, which ZeroMark calls a “fire control system,” has two components: a small computer that has sensors, like lidar and electro-optical, and a motorized buttstock.

a16z-backed ZeroMark wants to give soldiers guns that don’t miss against drones

The RAW Dating App aims to shake up the dating scheme by shedding the fake, TikTok-ified, heavily filtered photos and replacing them with a more genuine, unvarnished experience. The app…

Pitch Deck Teardown: RAW Dating App’s $3M angel deck

Yes, we’re calling it “ThreadsDeck” now. At least that’s the tag many are using to describe the new user interface for Instagram’s X competitor, Threads, which resembles the column-based format…

‘ThreadsDeck’ arrived just in time for the Trump verdict

Japanese crypto exchange DMM Bitcoin confirmed on Friday that it had been the victim of a hack resulting in the theft of 4,502.9 bitcoin, or about $305 million.  According to…

Hackers steal $305M from DMM Bitcoin crypto exchange

This is not a drill! Today marks the final day to secure your early-bird tickets for TechCrunch Disrupt 2024 at a significantly reduced rate. At midnight tonight, May 31, ticket…

Disrupt 2024 early-bird prices end at midnight

Instagram is testing a way for creators to experiment with reels without committing to having them displayed on their profiles, giving the social network a possible edge over TikTok and…

Instagram tests ‘trial reels’ that don’t display to a creator’s followers

U.S. federal regulators have requested more information from Zoox, Amazon’s self-driving unit, as part of an investigation into rear-end crash risks posed by unexpected braking. The National Highway Traffic Safety…

Feds tell Zoox to send more info about autonomous vehicles suddenly braking

You thought the hottest rap battle of the summer was between Kendrick Lamar and Drake. You were wrong. It’s between Canva and an enterprise CIO. At its Canva Create event…

Canva’s rap battle is part of a long legacy of Silicon Valley cringe

Voice cloning startup ElevenLabs introduced a new tool for users to generate sound effects through prompts today after announcing the project back in February.

ElevenLabs debuts AI-powered tool to generate sound effects

We caught up with Antler founder and CEO Magnus Grimeland about the startup scene in Asia, the current tech startup trends in the region and investment approaches during the rise…

VC firm Antler’s CEO says Asia presents ‘biggest opportunity’ in the world for growth

Temu is to face Europe’s strictest rules after being designated as a “very large online platform” under the Digital Services Act (DSA).

Chinese e-commerce marketplace Temu faces stricter EU rules as a ‘very large online platform’

Meta has been banned from launching features on Facebook and Instagram that would have collected data on voters in Spain using the social networks ahead of next month’s European Elections.…

Spain bans Meta from launching election features on Facebook, Instagram over privacy fears

Stripe, the world’s most valuable fintech startup, said on Friday that it will temporarily move to an invite-only model for new account sign-ups in India, calling the move “a tough…

Stripe curbs its India ambitions over regulatory situation

The 2024 election is likely to be the first in which faked audio and video of candidates is a serious factor. As campaigns warm up, voters should be aware: voice…

Voice cloning of political figures is still easy as pie

When Alex Ewing was a kid growing up in Purcell, Oklahoma, he knew how close he was to home based on which billboards he could see out the car window.…

OneScreen.ai brings startup ads to billboards and NYC’s subway

SpaceX’s massive Starship rocket could take to the skies for the fourth time on June 5, with the primary objective of evaluating the second stage’s reusable heat shield as the…

SpaceX sent Starship to orbit — the next launch will try to bring it back