Back to blog
Competition Hackathon

ShellHacks 2026

ShellHacks 2026, Florida International University

The JARVIS team in front of the ShellHacks wall Parker at the ShellHacks wall
The team, plus one of me at the wall. Big thanks to my teammates for a weekend of building, breaking, and fixing things together. Really happy to have been here.

ShellHacks is Florida's largest hackathon, a weekend at FIU of building something from nothing and demoing it to judges and sponsors on Sunday. I went in planning to build an AI voice clone detector, pivoted partway through the weekend, and came out with JARVIS: a desk robot you talk to that can see through its camera, turn to follow you, and actually do things on your computer when you ask.

View the code on GitHub

What JARVIS does

Most AI assistants live in a chat window and have no idea what's happening in the room. JARVIS sits on the desk instead. You talk to it out loud, and it can:

Here it is mid-demo. The webcam sits on the servo at the top of the tower with the Arduino wired in next to it, and the laptop shows what JARVIS sees: a green TRACKING box on the person in front of it, a blue skeleton over his hand from the gesture model, and no box at all on the guy walking past in the background. The status line at the bottom reads "Steadying the head (lag 450 ms)", the servo compensating for how late the camera image arrives.

JARVIS on the table: a webcam on a servo head, an Arduino, and the laptop showing the tracking view
JARVIS tracking one face and one hand, ignoring the room behind it.

How it's built

The hardware is about as scrappy as it gets. The tower is paint stir sticks and zip ties, the servo and webcam sit on a couple of small boxes on top, and the Surface laptop off to the side runs everything. Behind it was a trifold with my resume and a Blackstone careers flyer.

The full JARVIS table at ShellHacks 2026
The table at ShellHacks.

The software is where most of the work went:

PieceWhat it does
ElevenLabs Conversational AIThe voice. Every capability is a tool the agent decides to call on its own, 19 in total
GeminiThe eyes: answers questions about the camera and screen, finds objects by name, reads text
OpenCV + OpenCV ZooLocal models for faces (YuNet), hands and gestures (MediaPipe), and object tracking (VitTrack), no API calls
Arduino + servoThe head. Python sends angles over serial, the Arduino glides the servo there smoothly
PythonTies it together, with a live camera window, a control bar, keyboard shortcuts, and live captions

How it works

Every snippet below is trimmed from the real code in the JARVIS repo.

Talking turns into actions. Every Python function JARVIS can run is wrapped with one decorator. It registers the function by name, logs the call, and makes sure a failure comes back as a sentence the agent can say instead of a crash. A setup script pushes the matching list of 19 tool definitions to the ElevenLabs agent, so the model picks which one to call from what you said, no keyword matching.

def tool(fn):
    def wrapper(params):
        try: out = fn(params)
        except QuotaExhausted as e: out = str(e)
        except Exception as e: out = f"{fn.__name__} failed: {e}"
        return str(out)
    TOOLS[fn.__name__] = wrapper

@tool
def look(p):
    frame = current_view()   # current camera frame, including zoom
    ...                      # sent to Gemini with the user's question

Picking one face out of a crowd. YuNet finds every face in the frame, which at a hackathon is a lot of faces. JARVIS keeps only the ones that are big enough and at least 60% the size of the biggest face, so people walking by in the back get dropped. That's why the person walking past in the demo photo has no box.

big = out[0][3]   # height of the largest face
out = [b for b in out if b[3] >= FACE_MIN and b[3] >= big * 0.6]

Moving the head. Python only sends a target angle like 112,90 over serial. The Arduino does the motion itself, 66 times a second: it speeds up gently, cruises, and eases in as it lands, so the head never jerks. Python mirrors the same four numbers so it knows where the head is mid-turn.

float glide(float pos, int target, float &vel) {
  float d = target - pos;
  if (fabs(d) <= 0.5 && fabs(vel) <= ACCEL) { vel = 0; return target; }
  float want = constrain(d * EASE, -MAX_STEP, MAX_STEP);  // speed we'd like
  vel += constrain(want - vel, -ACCEL, ACCEL);            // ramp toward it
  return pos + vel;
}

Gestures. The MediaPipe hand models from OpenCV Zoo classify the hand, and each gesture maps to one action. A gesture only counts if it's held for about half a second and the hand belongs to the person JARVIS is tracking, so someone waving in the background can't take a photo.

GESTURE_ACTIONS = {
    "ThumbsUp":   ("photo in 3", start_photo_countdown),
    "Two":        ("record", toggle_recording),
    "Five":       ("mic", toggle_mic),
    "ThumbsDown": ("reset", reset_zoom_and_tracking),
}

Not talking to itself. The mic sits a foot from the laptop speakers. While JARVIS is speaking, and for a third of a second after, the audio going to the agent is swapped for silence.

if S.mic_muted or time.time() < S.speaking_until + 0.35:
    audio = b"\x00" * len(audio)   # send silence instead of the echo

Before and after

Both clips were recorded by JARVIS itself, from the camera on its head. The first is from before the tracking was dialed in. Watch the head overshoot, swing past the person in front of it, and swing back, with the whole frame smearing from motion blur.

Before: overshoot and hunting.

The second is the finished version. The person stays centered while the room slides past behind them, no overshoot, no hunting back and forth.

After: a smooth follow.

The difference came from one idea: stop assuming, and measure. On startup JARVIS nudges the head 8 degrees and watches where the face moves in the image. That tells it which way the servo is mounted, and how long the camera takes to show the move, which is the delay it has to lead by.

ok = (d > 0) == (PAN_DIR > 0)   # did the face move the way the model expects?
if not ok: PAN_DIR = -PAN_DIR    # servo is mounted the other way, flip it

What broke, and what I learned from it

Most of the weekend wasn't writing features, it was finding out why the real hardware didn't behave like the plan.

It said the word instead of doing the thing. The first time I asked what it could see, the agent just said "look" out loud. The tool existed but was never properly attached to the agent, so the model treated it as a word. I ended up scripting the entire agent setup so that couldn't drift again.

It talked to itself. The mic picked up JARVIS's own voice from the speakers, so it kept answering itself in a loop. Muting the mic while it speaks, plus a mute toggle, fixed it.

The head ran away. At one point the camera swung all the way to one side and stared at the ceiling. The servo was mounted the opposite way from what the code assumed, so every correction pushed the target further off center. The startup nudge above is the fix.

It wobbled. That's the first clip. The webcam image was arriving later than I'd assumed, so it kept reacting to where the person was a moment ago. Measuring the delay and slowing the servo down with a gentle ramp got it to the second clip.

It chased strangers. In a room full of hackers, the tracker kept jumping to people walking behind me. It now locks onto one main person, ignores small faces in the background, and holds the lock if I briefly turn away.

It ran out of eyes. The free Gemini tier allowed about 20 requests a day, and my background awareness loop burned through that fast. That pushed me to move everything I could onto local models, and to only call the cloud when a question actually needs it.

The biggest lesson of the weekend: the gap between "it works in my head" and "it works on the table" is where all the real engineering happened. Every fix that mattered came from watching the hardware misbehave and measuring why.

The people

Every photo below was taken by JARVIS, which is why they're all from the same low angle at webcam quality. Most were a thumbs up at the camera and a 3 second countdown. I got to show it to Waymo and to two people from Blackstone, and it was a good reminder that a working demo starts better conversations than any pitch does.

Visitor giving JARVIS a thumbs up to trigger a photo
A thumbs up to the camera, and JARVIS counts down and takes the shot
Talking through JARVIS with a visitor at the table
Walking a visitor through how it works
A Blackstone official at the JARVIS table
One of the Blackstone team
A second Blackstone official at the JARVIS table
And the second from Blackstone

A friend from the FAU Cybersecurity Club stopped by to check it out too, and the team got in front of it one more time.

The team, photographed by JARVIS
The team again, this time shot by JARVIS
A friend from the club giving two thumbs up to JARVIS
Two thumbs up, one photo

Key takeaway

I came in as a security person and left having built a robot, a voice agent, and a computer vision pipeline in one weekend. I used Claude as a coding partner to move faster, but the part I'm proudest of is the debugging: figuring out why a servo runs away, why a camera lies about timing, and how to make an assistant useful without letting it talk over you. The code is on GitHub.

Walking out at night, with the FIU sign lit up blue against the dark, it hit me how much fun the whole weekend was. Tired, a little sleep deprived, and really glad I was there. Thanks to the ShellHacks organizers and every sponsor who stopped by the table.

The FIU sign lit up blue on a campus building at night
FIU at night, on the way out.

View the repo on GitHub