Stepping back a moment, let's consider the history of instrumentation of all sorts. In the beginning, there was no instrument automation of any sort, just an experimenter manipulating the environment to drive the desired effect. For example, the original PCR was run by moving samples between water baths - and adding fresh polymerase each cycle!. The invention of thermocyclers automated this process, and introduced the question of how to set a thermocycling profile. Once you had profiles, then people would want to save the profiles. Luckily, PCs were getting both relatively powerful and also inexpensive in that same time period, so a computer with sufficient power could either be connected to a thermocycler or embedded within it.
And so the question of managing these protocols arose. Probably around the same time Windows started enabling longer file names - otherwise every protocol would have only eight character names! Macs also sometimes show up as lab computers, but my impression is PCs dominated. Really ancient machines might have even older computers - during my internship in the late 1980s there was a gamma counter which had an ancient Commodore PET computer attached, with the built-in screen burnt in to near illegibility.
Then came thermocycler mounted on liquid handing robot decks, thermocyclers as components of integrated workcells, and now thermocyclers in autonomous laboratories. Remember, thermocycler here is a proxy for all sorts of other instruments with complex protocols.
With all of these, there was a risk - what if I create a protocol and somebody else later edits it? What if I call my protocol "best_thermocycle_profile_ever" and somebody else foolishly claims the same name despite their different protocol being inferior to mine? How do I keep all the profiles synchronized? Our Nebula installation has over a dozen ATC thermocyclers and four Opus qPCR thermocyclers.
The simplest solution to ensure robustness is to not use named protocols. If my complete Catalyst workflow has the thermocycler profile built in to it, then I need not worry about someone editing anything on the instruments nor name collisions nor about whether the protocol lives on a given instrument. Thermocycler protocols can be described compactly, so it shouldn't be a bandwidth issue to just push them to the instrument at run time.
But the first potential pitfall is whether the instrument supports such a mode. What if the instrument maker doesn't enable pushing protocols at run time? Or do they allow it, but I still must name it? Or perhaps the device drivers I'm working through don't support the push - that's something I can at least pester a local software team to work on.
But that brings up more questions. If I can write a protocol, how do I know it is valid? One way is to make the protocol upload strongly constrained - say just a trio of temperature+time pairs for ordinary thermocycling - but that might unduly minimize the scope of protocols I can validate. Ideally, part of the instrument API (or a standalone tool) would include "validate this protocol". Similarly, it would be useful (or in our case, critical) to be able to ask the instrument to compute the runtime of a given protocol. Sure, you can try to estimate it from the cycle times, but there's always a risk of missing some little detail of the timing such as pre- and post- steps. That even gets complicated, as the agent told me in one of my API evaluations - do you want the timing for when the labware can be removed from the instrument or when data can be pulled off the system? For some instruments, those times might be very different.
Suppose we decide that files must live on the instruments - rearing that ugly issue of keeping a fleet in sync. Ideally I could ask the instrument to give me a cryptographic hash of the protocol, which would protect me from unobserved changes or protocol name reuse. A fallback would be being able to pull the file and compute the hash myself (well, really my protocol program).
But is that file some proprietary binary format, or something my coding agent can interpret? The latter is clearly far more desirable, as then my agent can browse a library of protocols on an instrument and see if one fits my needs. That would be particularly true if a protocol is parameterizable. Plus recording the exact protocol used is a key aspect of capturing every nuance of an experiment that autonomous labs promise.
By my very limited & naive survey, rigorous protocol versioning isn't a norm in the lab instrumentation world - actually, I'm hedging but expect the answer is "is rarer than hen's teeth". It's always been a risk, but higher degrees of automation raise the stakes considerably.
If you're an upstart automation company, I hope you'll read this and think "this isn't so hard to do better than the incumbents!". As usual, I worry that so many existing instruments are in the stables of very large players that want to spend as little investment on upgrading them as they can get away with. New, or relatively new, entrants looking for an edge can plunge straight to supporting highly robust, highly repeatable operations - a plus in any environment from personal device to shared autonomous laboratory.
No comments:
Post a Comment