Client work · 2016 · for Eyrus

DCS

RFID site tracking, from the box to the cloud.

Eyrus tracks the people and equipment on a construction site with RFID. I wrote the firmware-level software on the box beside each antenna, remotely managed, with health checks and automatic updates, along with the API its readings go to and the app that runs the fleet.

Client
Eyrus
We built
Device Firmware, API & Rails Web App
With
RFID, Raspberry Pi, Ruby, Rails, Sidekiq, Redis, Monit
The Eyrus logo, Know it all, above an iPad showing a device dashboard, on navy.

The brief

RFID on a job site

Eyrus tracks the movement of people and equipment at construction sites using RFID. Workers wear tags, antennas pick them up, and a small computer beside each antenna sends what it reads to the cloud.

The challenge

Nobody to reboot it

The hardware, an enclosure with a Raspberry Pi, battery power, charging and RFID antennas, was designed, prototyped and manufactured by Never Stop Building. It sat on job sites with no one to look after it, so software had to make all of it work, cope with lost connectivity and sudden power loss, and update itself without ever failing.

What we made

The box, the API and the fleet

I wrote the firmware-level software that runs on each box, with health checks and automated updates managed remotely, that gathers the readings from the RFID hardware, then the API that takes them in and the Rails app that configures and monitors the whole fleet, and passed the data on to Eyrus's .NET software.

A construction worker in a hard hat and safety vest viewing a device dashboard on an iPad at a building site
Illustration of a construction site with a crane, a building, and a worker walking between two containers fitted with receivers

The idea

A tag, an antenna and a box.

Workers on a site wear RFID tags, each one unique, so the system knows the moment someone walks past a receiver. Receivers go in strategic, identifiable places around the site, and they pick up each passing tag with a timestamp, the tag’s identifier, its battery level and more, and hand it to the nearest DCS box.

The box is a Raspberry Pi in a custom enclosure, with its own battery and charging. It does very little thinking on its own: it takes the raw reads from the antennas and posts them to my API, and everything after that, the processing and the management, happens in the cloud.

Diagram of a path through hallways past RFID antennas, with DCS 1 and DCS 2 units in the rooms beside it

Movement

Following someone through a site.

With antennas at the gates and doorways, and a tag on every worker, a person’s path through a job site shows up as a series of reads, one box after another. Those reads went up to the DCS site manager, a proxy that passed them on into Eyrus’s .NET ingestion point, where the analysis happened.

People were the main thing tracked: who was on a job site, for job costing, billing and payroll. Trucks and equipment carried tags too, so Eyrus could track deliveries coming and going, and watch equipment to detect and deter theft.

There were several antennas on a site, and everything turned on following a tag’s identifier from one to the next. That analysis lived in the .NET product. I was responsible for getting the raw reads off the antennas connected to a DCS box and into the proxy in an authenticated, safe way, and the proxy turned around and passed them on to be ingested.

The collection

Knowing when someone is really there.

The collection algorithm decides who is on a site from the antennas’ raw reads. When a worker enters a read zone, the tag is added to a list. While it is read, again and again, it stays on the list.

When the worker leaves, the reader waits for a set time, the persist time, in case the tag comes back. If it isn’t read again in that time, it comes off the list. If the worker wanders out and back in before the persist time runs out, the tag is never removed. The drawing below follows two workers, one walking straight through two zones and one wandering in and out.

Diagram of two read zones with an antenna in each. A worker walking straight through is added to the tag list when entering zone 1, stays on it while being read, comes off after the persist time, and is added and removed again for zone 2. A second worker wandering in and out is added, removed when the persist time expires, re-added on re-entering, and on the second zone re-enters before the persist time expires and is never removed.
DCS control circuit boards and power supplies laid out on a wooden workbench
A row of white Eyrus enclosures seen at a low angle along a workshop shelf, each with a black cable gland and the "Know it all." logoTwo installed Eyrus units with flat antennas, one on plywood inside a building under construction, one on a chain-link fence

The box

Hardware from someone else, software from me.

The hardware design, the prototyping and, in the end, the manufacturing were handled by Never Stop Building, and the person behind it was Jason Fox. That covers the enclosure, the Raspberry Pi and the choice of components. My part was everything that ran on it, and it was firmware in all but name: launch scripts, daemons and Sidekiq workers that start on their own when the box powers up, talk to the RFID hardware and the other physical pieces to gather their readings, watch the battery and power over the Pi’s GPIO pins, and report in. The cloud managed all of it remotely, with health checks and automated software updates.

It was an Internet of Things system, a fleet of small computers run from the cloud, and it had to work with nobody standing next to it. A box on a fence at a building site loses its network and loses its power, and nobody is there to sort it out.

Ten DCS control boards laid out on a workbench, each a Raspberry Pi with a small screen above a bank of blue relays on a red baseplate, beside a parts drawer and a hot-air rework station
The front of a finished DCS enclosure: the Eyrus logo, a three-position switch reading Enable, On and Standby, and three lights labelled Stat, Scan and Data
A DCS enclosure bolted to a pole beside a chain-link fence on a building site, with a grey junction box on a post next to it

Resilience

What a box does when things go wrong.

Nearly everything the box does goes through a queue. A daemon that reads a tag doesn’t call the API; it puts the read in Redis and moves on, and Sidekiq delivers it when it can. Monit watches the services and starts them in the order they depend on each other, so a box that loses power is meant to come back on its own.

Step through a day in the life of one box here. It’s a drawing of how it was designed to behave, not a recording.

DCS boxOnline

  • RedisRunning
  • SidekiqDelivering
  • HeartbeatRunning
  • RFID readerReading
  • Battery & powerOn mains

Tag readRedisSidekiqIngestion API

A read A worker walks past an antenna. The read goes into Redis first, and Sidekiq delivers it to the API a moment later.

DCS boxNo network

  • RedisHolding reads
  • SidekiqWaiting to deliver
  • HeartbeatCan’t reach the API
  • RFID readerStill reading
  • Battery & powerOn mains

Tag readTag readTag readIngestion API

Offline The connection drops. The reader only ever writes to Redis, so the reads wait there.

DCS boxOnline

  • RedisEmptying
  • SidekiqDelivering the backlog
  • HeartbeatReporting again
  • RFID readerReading
  • Battery & powerOn mains

Tag readRedisSidekiqIngestion API

Back online The network returns and Sidekiq delivers what was waiting, then carries on as before.

DCS boxOn battery

  • RedisRunning
  • SidekiqDelivering
  • HeartbeatRunning
  • RFID readerReading
  • Battery & powerWatching the voltage

Mains lostBatteryShut down below 11.7 V

Mains off The box runs on its battery. A daemon reads the battery’s voltage, and if it falls below 11.7 volts the box shuts itself down rather than dying.

DCS boxStarting

  • Redis1. Started first
  • Sidekiq2. After Redis
  • Heartbeat3. Starting
  • RFID reader4. Starting
  • Battery & power5. Starting

BootMonitEvery service, in order

Power back The box boots and starts monit, and monit starts the services in order, Redis first, so nothing runs without what it depends on.

DCS boxUpdating

  1. Stop the services
  2. Check out the release’s git tag
  3. Run the scripts this box hasn’t run
  4. Reload monit’s configuration
  5. Start the services
  6. Tell the API the update is done

CloudUpdateRoll back, if it fails

Update The box updates itself from a git tag, in the steps above, and the cloud can roll it back if the update doesn’t take.
Diagram of a cycle between the cloud Device Manager and a DCS 1 unit, with icons for updates, metrics, battery and status

The hard part

An update that can’t fail.

The hardest thing to get right was updating the software on boxes I couldn’t touch. What runs on a box is Ruby, not a compiled program, so an update means fetching a new set of code that has to start cleanly on a machine I may never see again, and a failed update means a box someone has to go and fix.

I engineered it out of git repositories and bash scripts, handing out versions as git tags on GitHub. A release is a git tag. The box stops its services, checks out the tag, runs any scripts it hasn’t run before, reloads monit’s configuration, starts everything again and tells the API it’s done. Those scripts were inspired by Rails migrations, for the one-off changes a box needs when it moves from one version to the next.

The update ran in isolation, with rollback from the cloud, and it had to survive everything that comes with a Ruby project on a small computer, Bundler included.

“I was working on a problem I hadn’t solved anything like before, and I still haven’t since. The RFID side was exciting too, and I had a lot of creative freedom to come up with elegant solutions to a complex problem.”

Matt, Lotus
The Device Manager's Device Details page for one box: green tiles for its heartbeat, software version and reporting, and a connectivity timeline that is green from 8:00 to 12:30.
Healthy every tile green, the timeline unbroken
The same page with trouble: a yellow heartbeat, a red software tile, a red turn-off tile reading Waiting, a yellow reboot tile reading Rebooting, and red and yellow breaks in the connectivity timeline.
In trouble the box waiting to be turned off, or rebooting, and gaps in its timeline

The fleet

Every box, and how it’s doing.

The cloud side was two separate pieces with different jobs. The first is the Rails management app and API, where Eyrus configures sites, gates, antennas, tags and the boxes themselves. The boxes ask it now and then for their settings and for reboot commands.

It was clear early on that detailed metrics and status would matter as much as the features, so each box reports its health and its heartbeat, and the app shows outages, restarts, the software version it runs and a connectivity timeline. From there a box can be restarted, or its software updated, from anywhere.

The ingestion API

Answer fast, deliver later.

The second piece is the ingestion API, which needs high uptime and low latency. It writes each event into Redis and answers straight away, and a separate Sidekiq process then forwards the events to Eyrus’s .NET software, which sat outside my cloud.

In July 2016 I ran load tests against it with siege, sending a heartbeat and event payload for two boxes at once, sized for a morning when about 500 people arrive at a site. These are the five runs, exactly as I recorded them.

Load tests against the ingestion API, July 2016
RunConcurrentLengthEventsAverageAvailable
1251 min5906.8 ms100%
2505 min5,9923.5 ms100%
31005 min11,4187.4 ms99.86%
41005 min11,580112 ms99.8%
510010 min21,499137 ms99.7%