<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://junruren.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://junruren.com/" rel="alternate" type="text/html" /><updated>2026-09-03T02:46:25+00:00</updated><id>https://junruren.com/feed.xml</id><title type="html">Junru Ren</title><subtitle>AI, autonomy, simulation, and product-building.</subtitle><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><entry><title type="html">A New Way Forward</title><link href="https://junruren.com/posts/2026/08/a-new-way-forward/" rel="alternate" type="text/html" title="A New Way Forward" /><published>2026-08-20T00:00:00+00:00</published><updated>2026-08-20T00:00:00+00:00</updated><id>https://junruren.com/posts/2026/08/A-New-Way-Forward</id><content type="html" xml:base="https://junruren.com/posts/2026/08/a-new-way-forward/"><![CDATA[<h2 id="pickup">Pickup</h2>

<p>On a Friday afternoon in August, a Waymo pulled up for me. Ahead of us: an hour of Bay Area rush-hour traffic, and my first autonomous ride on a freeway. Somewhere on the US-101, I realized I had spent the whole ride talking into my phone, dictating the first draft of this post. I would not have done that with a driver up front. This car had no one to overhear me.</p>

<p class="notice">This post also lives on <a href="https://junruren.substack.com/p/a-new-way-forward?utm_source=junruren.com&amp;utm_medium=referral&amp;utm_campaign=a-new-way-forward">my Substack</a> — comment there, or subscribe to get future posts by email.</p>

<p>By the time this reaches you, I will have been at Waymo for about a month as a product manager working on simulation. It is my first time as a product manager, after six years as a software engineer before MIT, and my first time working on autonomous vehicles, or, to use the current phrase, physical AI. One month in, I am taking on two firsts, with a car driving me around as I write about them.</p>

<figure>
  <a href="/images/2026-08-20-A-New-Way-Forward/Waymo-at-WaymoHQ.jpeg">
    <img src="/images/2026-08-20-A-New-Way-Forward/Waymo-at-WaymoHQ.jpeg" alt="The author smiling behind the open door of a white Waymo Jaguar I-PACE robotaxi parked on a sunny, tree-lined plaza at Waymo HQ, its rooftop lidar dome and sensor pods visible and a colorful mural painted on the open door" width="960" height="1280" />
  </a>
  <figcaption>August 2026: a Waymo picking me up from Waymo. The novelty has not worn off.</figcaption>
</figure>

<p>Waymo’s name stands for “a new way forward in mobility.”<sup id="fnref:1" role="doc-noteref"><a href="#fn:1" class="footnote" rel="footnote">1</a></sup> I borrowed it for the title because it fits me too. I have embarked on a new way forward in the kind of work I do and the industry I do it in. Both changes connect to a long-running wish to make the world a little better through something physical and tangible. (The irony is that simulation sounds like the total opposite of tangible. But hey, I believe a good simulator sits on the critical path of AI succeeding in our physical world.)</p>

<h2 id="worlds-not-just-words">Worlds, Not Just Words</h2>

<p>I believe the next decade of AI is about worlds, not just words: systems that understand space and physics, not only sentences. I think this matters most in the physical world, where a mistake costs more than a bad paragraph. Among the many bets on physical AI, autonomous vehicles look to me like the mature end of the frontier. They already carry ordinary people who pay for rides, and they have the data flywheel most embodied AI is still trying to start.</p>

<p>The hard question in autonomous driving is no longer whether a car can drive, but whether we can prove at scale that it drives well in situations nobody has seen yet. That is a simulation question. I came to this view in roughly the order I tell it here. The views are mine, worked out before I had a badge.</p>

<h2 id="why-i-care">Why I Care</h2>

<p>Self-driving cars were science fiction to me for most of my life. When Waymo was still a Google project,<sup id="fnref:2" role="doc-noteref"><a href="#fn:2" class="footnote" rel="footnote">2</a></sup> I filed it under admirable moonshots. It did not feel real to me until I sat in one.</p>

<p>Before I could sit in one, I learned how easy it is to take driving for granted. Between 2020 and 2022, I went through a series of injuries and surgeries. During stretches of those three years, I could not drive. In California, that is humbling. Every trip to a doctor, a grocery store, or a friend’s place became a small logistics problem. I am fine now, thankfully, but the experience left me with a lasting awareness of what driving actually requires.</p>

<figure>
  <a href="/images/2026-08-20-A-New-Way-Forward/Shopping-Cart.jpeg">
    <img src="/images/2026-08-20-A-New-Way-Forward/Shopping-Cart.jpeg" alt="A motorized mobility shopping cart with a wire basket, black seat, and tall orange safety flag parked in a wet grocery-store parking lot under a partly cloudy sky, with a sign on the basket reading 'IN-STORE USE ONLY'" width="960" height="1280" loading="lazy" />
  </a>
  <figcaption>Between surgeries, this was the only thing I could drive.</figcaption>
</figure>

<p>Driving requires a license, and many people, in California and around the world, do not have one for many reasons. It also requires a body that can drive, which mine temporarily could not.<sup id="fnref:3" role="doc-noteref"><a href="#fn:3" class="footnote" rel="footnote">3</a></sup> When you cannot drive and call a ride, there is a stranger in the car with you. Usually that is fine. Sometimes you would simply like the ride to be yours: to take a call, sit quietly, or get your hour back. Rush hour on the 101 while dictating a blog post turns out to be one of those times.</p>

<p>My first Waymo ride was on January 27, 2024, in San Francisco, when the service was already fully driverless but still gated behind a waitlist.<sup id="fnref:4" role="doc-noteref"><a href="#fn:4" class="footnote" rel="footnote">4</a></sup> A friend from college who worked at Waymo sent me an invite. I had just come out of the last of the surgeries, which is probably why I remember the ride the way I do.</p>

<p>On a single-lane street, a double-parked car blocked us. The Waymo waited, then pulled left into the oncoming lane, went around, and merged back. Two things happened in my head at once. First, that was exactly what a decent human driver would do. Second, I trusted it, because the screen in front of me showed what the lidar saw: nobody was coming. It was a small, educated, human-like decision, and it did more to make autonomy real to me than a decade of headlines had.</p>

<p>I was applying to MIT’s LGO program at the time. In an update I sent the program that spring, I wrote that the ride had shown me operational excellence outside of work: safety, deployment, fleet management, demand, regulation, all the work around the algorithm.</p>

<h2 id="words">Words</h2>

<p>I spent six years at SoundHound building voice AI: systems that turned what people said into what they meant, in cars and, later, on restaurant phone lines. They understood sentences, and I was proud of them. Then, in late 2022, a chatbot made much of what I knew feel less durable, which is roughly why I went back to school; I told that story in <a href="/posts/2026/07/mit-reflection/">the MIT post</a>.</p>

<p>What I did not expect was how quickly language would stop being the whole story. Language models are astonishing with words and oddly helpless with the world. Fei-Fei Li points out that even state-of-the-art multimodal models “rarely perform better than chance” at estimating distance, orientation, and size, or at mentally rotating an object.<sup id="fnref:5" role="doc-noteref"><a href="#fn:5" class="footnote" rel="footnote">5</a></sup> I recognized the shape of that gap from my own work. My systems could parse “turn left at the next light.” They had no idea what a light was, or a left.</p>

<p>The systems I had built understood sentences. The ones I wanted to work on next had to understand a street.</p>

<figure>
  <a href="/images/2026-08-20-A-New-Way-Forward/Words-Worlds-Wheels.jpeg">
    <img src="/images/2026-08-20-A-New-Way-Forward/Words-Worlds-Wheels.jpeg" alt="Hand-drawn sketch on aged notebook paper: a speech bubble filled with scribbled text, an arrow pointing to a wireframe cube containing a tree, a winding road, and a stick-figure pedestrian, and a second arrow pointing to a small rounded driverless car with a rooftop sensor, drawn speeding away" width="1600" height="893" loading="lazy" />
  </a>
  <figcaption>The argument, in one sketch: words, then worlds, then wheels.</figcaption>
</figure>

<h2 id="worlds">Worlds</h2>

<p>At MIT I went looking for the other modalities. In spring 2025 I took an advanced computer-vision class, and one project asked us to reconstruct a bulldozer in three dimensions from a single photo: given one angle, produce the front, the back, the side. Those classes and projects also redrew my map of the field. The research frontier was no longer one specialized model per perception task, an object detector here, a segmentation model there. Attention had moved to general models that learn a representation of the world across modalities and use it to predict what happens next. A few months later, the whole industry had a name for what I was circling: world models, systems that do not just recognize a scene but can imagine how it continues. Li was arguing in <a href="https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence">“From Words to Worlds”</a> that spatial intelligence is AI’s next frontier: “Without spatial intelligence, AI is disconnected from the physical reality it seeks to understand and cannot effectively drive our cars or guide robots.” I did not need much convincing.</p>

<p>The obvious applications are in the physical world. Physical AI covers humanoid robots, warehouse arms, and general-purpose robot policies. Much of it is dazzling, but by its own researchers’ account it is short of the one thing language models had in abundance. The largest open robot-learning dataset holds about a million trajectories across 22 kinds of robot; language models train on trillions of tokens.<sup id="fnref:6" role="doc-noteref"><a href="#fn:6" class="footnote" rel="footnote">6</a></sup> Most of that work is still in the lab. I knew I wanted to work at the frontier and in the real world at the same time.</p>

<h2 id="wheels">Wheels</h2>

<p>The exception I found was autonomous vehicles. They are robots with four wheels navigating the physical world, and they were already on the road: paying riders in roughly fifteen U.S. metros,<sup id="fnref:7" role="doc-noteref"><a href="#fn:7" class="footnote" rel="footnote">7</a></sup> and more than 200 million fully autonomous miles by Waymo’s public count.<sup id="fnref:8" role="doc-noteref"><a href="#fn:8" class="footnote" rel="footnote">8</a></sup> Every one of those miles produces data no lab could have invented from scratch. That data feeds model development and deployment, which lead to more miles. This is the flywheel most embodied AI is still trying to start. It made autonomous driving look to me like the mature end of the frontier. The work is not finished, but it is out of the lab.</p>

<p>Waymo also became one of MIT LGO’s industry partner companies in 2023,<sup id="fnref:9" role="doc-noteref"><a href="#fn:9" class="footnote" rel="footnote">9</a></sup> around the time I applied to the program. That gave me reasons to pay closer attention: partner introductions, classmates and alumni who worked there, and a lot of homework on my side so that I would not ask questions a search engine could have answered. My curiosity went first to the driver, the AI behind the wheel. But two years of an operations program did their work on me. I grew a real respect for everything that has to be true for a robotaxi to be safe, available, and legal in a city, and I would have happily worked on that side too. By the time full-time recruiting came around in late 2025, Waymo was at the top of my list.</p>

<p>There is a confession about temperament here. At SoundHound, I worked on in-car voice assistants for years before I finally sat in a car in China, in 2024, that ran the Mandarin voice system I had helped build. It took more than five years to close that loop. I am, in a way, an impatient person. I like to see my work reach real people sooner. Waymo’s driver was already on the road, and that mattered to me more than I expected.</p>

<h2 id="the-question-that-moved">The Question That Moved</h2>

<p>People disagree about whether autonomous driving is still a scientific problem or now mainly an engineering one. Andrej Karpathy, who led AI at Tesla for five years, described the remaining work this way last October: “this is not even near done… it’s a march of nines. Every single nine is a constant amount of work.”<sup id="fnref:10" role="doc-noteref"><a href="#fn:10" class="footnote" rel="footnote">10</a></sup> Waymo’s own co-CEO says the demo took 18 months and the product took about 15 years.<sup id="fnref:11" role="doc-noteref"><a href="#fn:11" class="footnote" rel="footnote">11</a></sup></p>

<figure>
  <a href="/images/2026-08-20-A-New-Way-Forward/DD-at-YC.jpeg">
    <img src="/images/2026-08-20-A-New-Way-Forward/DD-at-YC.jpeg" alt="Conference-hall view of YC Startup School 2026: Dmitri Dolgov speaks on a round orange stage beneath a large screen showing a slide titled '7 Lessons - Learned from shipping the most mature manifestation of AI in the physical world,' with three generations of Waymo vehicles pictured" width="960" height="1280" loading="lazy" />
  </a>
  <figcaption>One week into the job, at YC Startup School: Dmitri Dolgov's seven lessons. Several of them ended up in this post's footnotes, and his slide title is probably where "the mature end of the frontier" started forming in my head.</figcaption>
</figure>

<p>I think the scientific question moved rather than disappeared. It is no longer “can a car drive?” but “can we prove, at scale, that it drives well in situations nobody has seen yet?” Answering that depends on simulation. A simulator has to reproduce the world realistically, generate rare events that real roads seldom offer, and evaluate a driver against them before it meets them. Waymo has been public about this. Simulation is one of the three pillars of its approach to demonstrably safe AI, and its newest simulator is itself a world model, generating camera and lidar data for events “from a tornado to a casual encounter with an elephant.”<sup id="fnref:12" role="doc-noteref"><a href="#fn:12" class="footnote" rel="footnote">12</a></sup> For me, the words-to-worlds argument leads from the car to the simulator inside the car company.</p>

<p>My new title, product manager for simulation realism, is essentially the sim-to-real gap with a job description attached. A simulator is also where the frontier lands: world models come out of research upstream, and a simulator turns them into a working product, one that trains and validates a robotic driver. Standing at that junction takes someone who can read the papers and ship a product. Two years at MIT rebuilt the first muscle; six years of building software gave me the second. For me it looks like a sweet spot, though one month in, that is still a working hypothesis of mine. The role touches the whole system: the data coming off the fleet, the models, what has to be true before something ships, and the streets it all has to survive. Simulators, evaluation, and closing the gap between the simulated and the real are problems every embodied AI will need to solve; <a href="https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence">Li’s essay says as much about robots</a>. The ideas travel. That is why this feels like work I could stay close to for a long time.</p>

<p>I hope robot-learning data catches up soon, and autonomous vehicles stop being the exception in this argument. That would be good news for everyone, and the ideas would only travel further.</p>

<h2 id="the-deputy-pm">The Deputy PM</h2>

<p>The other first needs more explaining, because I did not set out to become a product manager. I came to MIT mainly for a technical reset.</p>

<p>I did not see the pattern clearly at the time. As a junior engineer at SoundHound, I felt comfortable huddling with product and program managers, mostly because I wanted to write code that would still make sense in a year. As a tech lead later on, I found myself bringing intuitions about what a feature should look like and how to spec it, sitting across from my product-manager counterpart almost as an equal partner. I listened to real users’ voice interactions with our system. I flew to Nashville, where our sales and support teams sat, to hear what restaurant owners were actually asking for. Colleagues from sales and customer success started coming to me with product questions, and I enjoyed explaining the same piece of code three different ways for three different audiences. Looking back, that was a deputy product manager without the label. At the time I would have said I was just an engineer who wanted the bigger picture.</p>

<p>Product management did not feel like a plausible next step until I read product-manager job descriptions closely for the first time. I started to recognize myself in them: a deep technical background, an understanding of AI systems, and the strategy and people skills I had been practicing without naming. I worked backwards from those descriptions to my own history, and I applied.</p>

<p>If the hard problems in physical AI cut across evaluation, operations, safety, and all the work around the algorithm that I noticed on that first ride, then a technically deep product-manager role seems like a reasonable place for me. That is the argument, anyway. I will report back on whether it survives contact with the job.</p>

<p>Some of the vocabulary makes me wince. Product managers now call themselves builders, and I can hear the MBA in myself when I say “stakeholder mapping.” Cringe aside, the substance is real. A technically deep PM with today’s AI tooling can prototype and test ideas fast, and the good old skills, understanding organizations, incentives, and people, are what turn a prototype into something a team will actually build. I still write code, happily, when it makes me faster. I have just stopped wanting to be measured by it. Among engineers I was probably the most PM-looking one. Among PMs, I hope to stay the one most comfortable thinking like an engineer.</p>

<h2 id="one-month-in">One Month In</h2>

<p>One month is too early for conclusions, and most of the specifics belong at work. I did my homework before starting, but it did not make the pivot smaller. I really am changing both what I do and the industry I do it in at the same time. For now, I need to learn from the inside: how things actually work here and where I can contribute most. Colleagues, new and long-tenured alike, tell me the same thing in different words: at Waymo, you learn something new every day. I cannot yet claim to know one percent of what a good product manager is. Then again, I should not be that conservative. I know more than I did a month ago.</p>

<p>Everyone I have met so far, regardless of seniority or background, has been welcoming, generous with their time, and patient with a career pivoter. I look forward to working with them and learning from them.</p>

<h2 id="drop-off">Drop-Off</h2>

<figure>
  <a href="/images/2026-08-20-A-New-Way-Forward/2024-First-Surface-Street-Ride-and-2026-First-Freeway-Ride.jpeg">
    <img src="/images/2026-08-20-A-New-Way-Forward/2024-First-Surface-Street-Ride-and-2026-First-Freeway-Ride.jpeg" alt="Two photos from inside a driverless Waymo with an empty driver's seat. Left: a dusk city street in San Francisco with a gold-domed building ahead, the screen reading 'Arrival in 14 min at 5:29 PM.' Right: a daytime freeway with the wheel turning itself, the screen reading 'Arriving in 55 min at 5:54 PM.'" width="1280" height="853" loading="lazy" />
  </a>
  <figcaption>Left: my first Waymo ride, San Francisco, January 27, 2024. Right: my first freeway ride, two and a half years later — the one this post was dictated in.</figcaption>
</figure>

<p>An hour after the pickup, the car took the exit toward my drop-off. Waymo has said publicly that the same driver could one day power trucks and personally owned vehicles,<sup id="fnref:13" role="doc-noteref"><a href="#fn:13" class="footnote" rel="footnote">13</a></sup> so the mission of getting more people moving is bigger than the robotaxi I was sitting in. For me, that mission stays personal. Three years ago I could not drive. Now I can be carried an hour across the Bay Area while I write, trusting a machine in a way I could not before. I want that for the many people who cannot drive, and for the many who can but would sometimes rather not. I also like that what I work on now is something my family and friends can see and ride: AI embodied on the street, in addition to a voice inside a dashboard.</p>

<p>The months of learning ahead will outnumber the years behind me. I hope they add up to a meaningful contribution to a mission that is personal to me. I will write more here about worlds and simulation as I learn.</p>

<p>Thank you to my mom and dad, who drove me around when I could not drive myself. I hope I can return the favor and drive them around now, even if the car does the driving.</p>

<!-- HOW-I-MADE-THIS statement:

Whenever I had a moment, I'd record myself sharing a snippet of stories or ideas, partly from the back seat of the Waymo :) Then, with my recordings transcribed, I used AI to help me tabulate all the stories, and then I'd organize them into a storyline. Often I'd ask AI to suggest some funny and witty ways to word something. I also used AI for some literature review and fact-checking. Finally, I used AI to check my grammar, as English is not my first language.
-->

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:1" role="doc-endnote">
      <p>From Waymo’s December 2016 announcement: <a href="https://medium.com/waymo/say-hello-to-waymo-whats-next-for-google-s-self-driving-car-project-b854578b24ee">“Waymo stands for a new way forward in mobility.”</a> <a href="#fnref:1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:2" role="doc-endnote">
      <p>The project began in January 2009 as the Google Self-Driving Car Project inside X. Fully driverless rides opened to the general public in the Phoenix area in <a href="https://waymo.com/blog/2020/10/waymo-is-opening-its-fully-driverless.html">October 2020</a>. <a href="#fnref:2" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:3" role="doc-endnote">
      <p>U.S. DOT Bureau of Transportation Statistics, <a href="https://www.bts.gov/newsroom/travel-patterns-american-adults-disabilities">“Travel Patterns of American Adults with Disabilities”</a>: an estimated 25.5 million Americans age 5 and older have self-reported travel-limiting disabilities, and about 3.6 million of them do not leave home at all (2017 National Household Travel Survey). <a href="#fnref:3" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:4" role="doc-endnote">
      <p>San Francisco timeline: paid, fully driverless rides were <a href="https://www.cpuc.ca.gov/news-and-updates/all-news/cpuc-approves-permits-for-cruise-and-waymo-to-charge-fares-for-passenger-service-in-sf-2023">approved in August 2023</a>; the waitlist came down and the service <a href="https://waymo.com/blog/2024/06/waymo-one-is-now-open-to-everyone-in-san-francisco/">opened to everyone in June 2024</a>. <a href="#fnref:4" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:5" role="doc-endnote">
      <p>Fei-Fei Li, <a href="https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence">“From Words to Worlds: Spatial Intelligence Is AI’s Next Frontier”</a>, November 2025. On robots: “World models will play a defining role in scaling robotic learning by closing the gap between simulation and reality.” <a href="#fnref:5" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:6" role="doc-endnote">
      <p><a href="https://arxiv.org/abs/2310.08864">Open X-Embodiment</a> pools about one million trajectories across 22 robot embodiments. A 2025 study of data scaling in robotic manipulation opens by noting that “the principles of effective data scaling in robotic manipulation remain insufficiently understood” (<a href="https://arxiv.org/abs/2507.06219">Shi et al.</a>). <a href="#fnref:6" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:7" role="doc-endnote">
      <p>Waymo counted <a href="https://waymo.com/blog/2026/02/dallas-houston-san-antonio-orlando/">ten commercial metro areas</a> in February 2026 and began fully autonomous driving in <a href="https://waymo.com/blog/shorts/ro-den-lv-sd-tmpa/">four more</a> that July. <a href="#fnref:7" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:8" role="doc-endnote">
      <p><a href="https://waymo.com/blog/2026/02/dallas-houston-san-antonio-orlando/">“Over 200 million fully autonomous miles traveled,”</a> Waymo, February 2026; Dolgov cited 220 million-plus in <a href="https://www.ycombinator.com/library/WV-waymo-co-ceo-dmitri-dolgov-the-demo-is-only-1-of-the-work">July 2026</a>. <a href="#fnref:8" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:9" role="doc-endnote">
      <p><a href="https://lgo.mit.edu/partner-companies/">LGO partner companies</a>: Waymo, partner since 2023. <a href="#fnref:9" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:10" role="doc-endnote">
      <p>Andrej Karpathy on the <a href="https://www.dwarkesh.com/p/andrej-karpathy">Dwarkesh Podcast</a>, October 2025. <a href="#fnref:10" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:11" role="doc-endnote">
      <p>Dmitri Dolgov, <a href="https://www.ycombinator.com/library/WV-waymo-co-ceo-dmitri-dolgov-the-demo-is-only-1-of-the-work">YC Startup School, July 2026</a>: “The demo took 18 months; the product took about 15 years.” <a href="#fnref:11" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:12" role="doc-endnote">
      <p>Waymo, <a href="https://waymo.com/blog/2026/02/the-waymo-world-model-a-new-frontier-for-autonomous-driving-simulation/">“The Waymo World Model: A New Frontier for Autonomous Driving Simulation”</a>, February 2026: “Simulation is a critical component of Waymo’s AI ecosystem and one of the three key pillars of our approach to demonstrably safe AI.” <a href="#fnref:12" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:13" role="doc-endnote">
      <p>Dolgov, YC Startup School, July 2026: “In the future, we’ll power different products and different commercial applications like trucking and personally owned vehicles.” <a href="#fnref:13" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><category term="personal" /><category term="waymo" /><category term="career" /><summary type="html"><![CDATA[I've joined Waymo as a product manager. Why I think the next decade of AI is about worlds, not just words, and why a robotaxi is where I chose to work on it.]]></summary></entry><entry><title type="html">One Day, Two Years at MIT</title><link href="https://junruren.com/posts/2026/07/mit-reflection/" rel="alternate" type="text/html" title="One Day, Two Years at MIT" /><published>2026-07-29T00:00:00+00:00</published><updated>2026-07-29T00:00:00+00:00</updated><id>https://junruren.com/posts/2026/07/mit-reflection</id><content type="html" xml:base="https://junruren.com/posts/2026/07/mit-reflection/"><![CDATA[<p>Two months ago, in May 2026, I somehow graduated from MIT four times.<sup id="fnref:1" role="doc-noteref"><a href="#fn:1" class="footnote" rel="footnote">1</a></sup> Somewhere in that run of ceremonies, in front of the Great Dome overlooking the Charles River, I listened to <a href="https://news.mit.edu/2026/commencement-address-lisa-su-0528">Dr. Lisa Su’s inspiring speech</a> about being ambitious, solving hard problems, and finding ways to make luck. I also snacked one last time on bananas from the Banana Lounge. The final result was two degrees, an S.M. in EECS and an MBA, through the <a href="https://lgo.mit.edu/">Leaders for Global Operations (LGO) program</a> with a $100,000 fellowship.</p>

<p class="notice">This post also lives on <a href="https://junruren.substack.com/p/mit-reflection?utm_source=junruren.com&amp;utm_medium=referral&amp;utm_campaign=mit-reflection">my Substack</a> — comment there, or subscribe to get future posts by email.</p>

<figure>
  <a href="/images/2026-07-29-MIT-Reflection/Dome.jpeg">
    <img src="/images/2026-07-29-MIT-Reflection/Dome.jpeg" alt="Junru in MIT graduation regalia holding two red diploma folders with two companions in front of the Great Dome" width="1280" height="853" />
  </a>
  <figcaption>School of Engineering &amp; Schwarzman College of Computing Advanced Degree Ceremony in front of the Great Dome, May 27, 2026.</figcaption>
</figure>

<p>Since then, I have sat down for about half an hour every other day and tried to relive the previous 24 months. I have recorded monologues, chatted with prospective students, reread old essays, and gone back through small stories that would never make a graduation speech. I want to write them down before I get preempted by the “ambitious problem solving” I am taking on after graduation.</p>

<p>This post is my attempt to give those two years a proper conclusion.</p>

<p>Years from now, I want this post to bring me back to the MIT I experienced. I am also writing for the classmates who went through these two years with me, assuming you are not already tired of one more blog post from me. At least this one is not another tutorial about <a href="/posts/2025/06/LaTeX-VSCode/">using LaTeX in VS Code</a> or <a href="/posts/2025/08/MIT-Thesis-LaTeX/">writing an MIT thesis</a>, or another set of <a href="/posts/2025/09/Cheatsheets/">cheat sheets</a>.</p>

<p>And finally, if you are interested in pursuing more school, especially mid-career, I hope you will find some interesting takes here.</p>

<h2 id="birds-eye-view">Bird’s-Eye View</h2>

<p><a href="https://lgo.mit.edu/">MIT’s Leaders for Global Operations (LGO) program</a> began in 1988 as Leaders for Manufacturing (LFM), at a time when Japan and other overseas rivals were challenging U.S. manufacturing dominance<sup id="fnref:2" role="doc-noteref"><a href="#fn:2" class="footnote" rel="footnote">2</a></sup>. The program has since evolved and expanded the range of students attracted to it. By the time I enrolled, the program appeared to be one of the most structured MS + MBA programs in the U.S., with a clearly defined timeline<sup id="fnref:3" role="doc-noteref"><a href="#fn:3" class="footnote" rel="footnote">3</a></sup>.</p>

<p>My LGO journey began in June 2024 with a highly structured summer semester attended only by the 45 students in the LGO Class of 2026. Over the summer, we quickly bonded while knocking out required classes that counted toward our dual-degree requirements. Then, in my first fall, I mixed the Sloan MBA core with engineering classes. I also became part of a roughly 70-student MBA cohort (Indian Ocean<sup id="fnref:4" role="doc-noteref"><a href="#fn:4" class="footnote" rel="footnote">4</a></sup>, let’s go!) as all LGO students integrated into the regular MBA program. January brought a three-week domestic plant trek. After the following spring, I left for a six-month internship and thesis period, then came back for one final spring at MIT. My schedule never stayed the same for very long.</p>

<p>Well. Nevertheless, allow me to fabricate a “typical day.” Below, I stitched together a bike ride from one semester, classes and lunches from several others, and a late-night search that happened on one very specific Monday. Every scene happened, just not on the same day.</p>

<h2 id="one-typical-day">One Typical Day</h2>

<h3 id="815-am---vassar-street">8:15 A.M. - Vassar Street</h3>

<p>During my first year, I lived on the far west side of campus. On mornings with an 8:30 class at Sloan, I usually left home with exactly enough time to bike there. I refused to skip breakfast, even after a late night. This meant I had only one remaining way to make up time: pedal faster.</p>

<p>I would get on my retro French road bike (my best investment at MIT) and ride east along Vassar Street. In winter, pristine snow covered the athletic field beside the road. The cold required the heaviest coat I had ever owned, a coat I had bought only after moving to Boston.</p>

<figure class="half">
  <a href="/images/2026-07-29-MIT-Reflection/Bike.jpeg">
    <img src="/images/2026-07-29-MIT-Reflection/Bike.jpeg" alt="A red Peugeot road bike leaning against a green railing beside the Charles River, with the Boston skyline and Harvard Bridge behind it" width="960" height="1280" loading="lazy" />
  </a>
  <a href="/images/2026-07-29-MIT-Reflection/Snow.jpeg">
    <img src="/images/2026-07-29-MIT-Reflection/Snow.jpeg" alt="MIT's red Briggs Field scoreboard seen through heavy snow, with campus towers fading into the background" width="1024" height="1280" loading="lazy" />
  </a>
  <figcaption>My red Peugeot road bike by the Charles, and Briggs Field in a snowstorm beside my route along Vassar Street.</figcaption>
</figure>

<p>Most of my early Sloan classes took attendance, sometimes through a code shown in the room. Missing the code meant losing one of the easiest points in the class, so the morning ride was rarely scenic, though I still tried my best to enjoy it.</p>

<p>The LGO class ahead of us had also passed down a favorite expression: GDM, or “grades don’t matter.” When I first heard it during Summer Core, my reaction was basically: yes, I understand, but I still plan to try very hard.</p>

<p>By my final spring, I had fully adjusted my mindset. I still cared about doing good work, sometimes too much, but I became more judicious about where I spent my time and more comfortable missing a few points to attend to other priorities. I started asking whether that hour belonged on the current assignment at all. Maybe it should go to another class, office hours, dinner with someone, or sleep.</p>

<h3 id="830-am---how-i-ended-up-in-an-mba-classroom">8:30 A.M. - How I Ended Up in an MBA Classroom</h3>

<p>The MBA part of my morning requires some explanation because I did not originally set out to earn one.</p>

<p>After six years in industry, I started by looking for a master’s program in computer science. ChatGPT’s public arrival in late 2022 made the search feel more urgent. I had spent my career building voice AI products, and suddenly some technical assumptions that had felt durable no longer did. I wanted time to study what was changing underneath the products, and I didn’t want to watch the field move from the same seat.</p>

<p>Before making a full commitment, I gave student life a trial run. In 2023, I took Stanford’s CS221 course through Stanford Online. “Online” did not mean I watched a few videos after work. For one quarter, I attended lectures on the Stanford campus, went to office hours, worked with classmates, and sometimes studied in the basement of the Huang Engineering Center. I liked being a student again.</p>

<p>Meanwhile, I went through the computer science graduate-program websites of the schools I admired. MIT looked like a dead end because the terminal EECS master’s route I had imagined was generally unavailable to external applicants. Farther down the page, though, I found one exception: LGO, a dual-degree program combining an engineering master’s degree with an MBA. The MBA was not at the top of my mind. At first, I mostly saw LGO as the route that made MIT possible.</p>

<p>And that is how I arrived at MIT after a very rigorous admissions journey.</p>

<figure>
  <a href="/images/2026-07-29-MIT-Reflection/Dome_2011_2024.jpeg">
    <img src="/images/2026-07-29-MIT-Reflection/Dome_2011_2024.jpeg" alt="Side-by-side photos of Junru standing in front of MIT's Great Dome in 2011 and 2024" width="1280" height="720" loading="lazy" />
  </a>
  <figcaption>From a campus visit in 2011 to arriving as an MIT student in 2024.</figcaption>
</figure>

<p>On the weekend before LGO began in 2024, I recreated a photo of myself taken in 2011 in front of the Great Dome. On LinkedIn, I <a href="https://lnkd.in/p/gHsvE9-X">wrote</a>:</p>

<blockquote>
  <p>The 16 years old boy touring the Massachusetts Institute of Technology campus in 2011 probably had no idea that one day he’d return to the exact spot as an MIT student…</p>
</blockquote>

<p>In fact, the 2024 me still had only a rough idea of what the non-CS half of MIT would do for me.</p>

<p>I cannot point to one class and say that it converted me into a business-school believer. Memo assignments changed how I organized an argument. Operations gave names to patterns I had never noticed in technical organizations. Reflecting on leadership made me look harder at what my decisions did to other people. Gradually, I stopped treating the MBA and the operations focus as extras attached to the MIT exception.</p>

<h3 id="1000-am---what-sloan-actually-taught-me">10:00 A.M. - What Sloan Actually Taught Me</h3>

<p>A Sloan class often began long before I entered the room. There might be a case to read, calculations to prepare, or a memo to submit. Even without a written assignment, a cold call from a professor could expose my weak preparation very quickly.</p>

<figure>
  <a href="/images/2026-07-29-MIT-Reflection/Sloan.jpeg">
    <img src="/images/2026-07-29-MIT-Reflection/Sloan.jpeg" alt="Junru wearing graduation cords and an LGO stole raises his hand in an empty Sloan classroom" width="1280" height="853" loading="lazy" />
  </a>
  <figcaption>One more in-class participation, this time in graduation mode.</figcaption>
</figure>

<p>I did not love every part of this system. Attendance points sometimes annoyed me, and participation incentives sometimes rewarded speaking for the sake of speaking. Still, those trivial “complaints” aside, business school gave me a lot of tools that I can use:</p>

<ul>
  <li>to communicate better (e.g., through writing executive memos in strategy classes and giving impromptu speeches),</li>
  <li>to understand myself better (e.g., through <a href="https://mitsloan.mit.edu/alumni/how-id-lab-creates-connections-and-busts-leadership-myths">ID Lab</a>),</li>
  <li>to understand others better (e.g., through Organizational Processes (OP)), and</li>
  <li>to basically poke into all sorts of random areas and explore broadly (e.g., I studied topics across finance, law, and blockchain, and I sat in the classrooms of Simon Johnson, a 2024 laureate of the Nobel Memorial Prize in Economic Sciences, and Gary Gensler, former chairman of the U.S. Securities and Exchange Commission).</li>
</ul>

<p>Operations classes as part of the LGO requirements also gave me concepts I could reuse elsewhere. The bullwhip effect made me look for the handoffs that turn a small change into a huge reaction. Pooling made me question why teams guarded separate resources that might work better together. Even a basic inventory tradeoff, the cost of too much versus the cost of too little, changed how I thought about decisions under uncertainty.</p>

<p>The examples came from factories, supply chains, and services, but I started noticing the same patterns in technology and research. This was when the “O” in LGO began to make sense to me.</p>

<h3 id="1145-am---lunch-coffee-chats-and-core-team">11:45 A.M. - Lunch, Coffee Chats, and Core Team</h3>

<p>The cleanest opening in many MIT calendars was between 11:30 and 1:00. That was when I scheduled coffee chats, met classmates, asked an upperclass student a question, or later talked with a first-year student trying to understand LGO.</p>

<p>“Coffee chat” can sound painfully transactional. Many of mine were just lunches with people I wanted to stay connected to. With everyone pulled toward different classes, recruiting schedules, clubs, and research, I had to maintain relationships with intention and sincerity. I think coffee-chat culture is a unique kind of magic in business school: it gives everyone a reason to meet people they might not otherwise get to know.</p>

<p>My LGO Summer team kept meeting for meals after the formal team structure ended. We were Team 6, so naturally we found a way to make a Six Sigma joke out of our name. By the end of the two years, we were still finding time for full-team dinners or lunches even though no class required us to.</p>

<p>LGO’s dedicated program office handled a good amount of boring but important coordination. It gave us a cohort, mapped a route through two schools, and supplied a skeleton for those demanding 24 months.</p>

<p>The skeleton still left plenty open. I chose my technical courses, looked for my own research home, and decided how much of the MBA social world I wanted.</p>

<p>I reduced some club commitments rather than keep titles without doing the work seriously. I skipped social events when an engineering problem set or exam had become the fixed priority. Other times I protected an experience because I knew I would not get another MIT semester later. I made some of these choices late and probably got a few of them wrong. Trying to do Sloan and EECS at full intensity was impossible, and I learned that late, but thankfully not too late.</p>

<h3 id="100-pm---eecs-and-research">1:00 P.M. - EECS and Research</h3>

<p>Afternoons often shifted from Sloan to EECS. The work changed quickly: fewer cases and memos, more code, equations on the chalkboard, experiments, and problem sets.</p>

<p>Over my two years, I took three classes that represented my core EECS interests:</p>

<ul>
  <li>6.7960 Deep Learning, taught by <a href="https://web.mit.edu/phillipi/">Phillip Isola</a>, <a href="https://beerys.github.io/">Sara Beery</a>, and <a href="https://jeremybernste.in/">Jeremy Bernstein</a>.</li>
  <li>6.8300 Advances in Computer Vision, taught by <a href="https://www.vincentsitzmann.com/">Vincent Sitzmann</a>.</li>
  <li>6.8610 Quantitative Methods for Natural Language Processing, taught by <a href="https://www.mit.edu/~jda/">Jacob Andreas</a>, <a href="https://omarkhattab.com/">Omar Khattab</a>, and <a href="https://www.chriswtanner.com/">Chris Tanner</a>.</li>
</ul>

<p>I chose them to survey the fundamentals of the current AI revolution across different modalities.</p>

<p>These three classes were structured similarly, with problem sets (math + coding) and a final research project. I had a rough start in my first class because I had not allocated enough time or energy to read religiously, think, and identify a research idea that was both novel and viable within the scope of half a semester. (Alas, I should have recalled a famous quote from my favorite undergraduate CS lecturer, <a href="https://cseweb.ucsd.edu/~ricko/">Rick Ord</a>: “Start early, start often.”)</p>

<p>Lesson learned. For a later computer-vision project, I started much earlier. I spent time understanding a small corner of the research frontier, read the future-work sections of papers, and checked what I could actually test with the available time, data, and compute. I also went to the teaching staff’s office hours while there was still time for their feedback to change the project. All of this sounds like common sense. When I was in the thick of the program, though, I needed some setbacks to realize it and refine my approach.</p>

<p>The project tested whether the order of objects in captions affected CLIP’s matching accuracy. I wrote up the technical details in <a href="/posts/2025/05/6.8300-final/">a separate post about the 6.8300 final project</a>. The part I remember best now is one short office-hour conversation.</p>

<p>Professor Phillip Isola was not teaching the class, but he held open office hours. I wanted to hear his thoughts because I appreciated his expertise in representation learning, so I showed up, explained the problem, and walked him through my experiment. In less than twenty minutes, he suggested a post hoc fix that could address the bias I had observed: permute the order of the candidate objects, run the model across those permutations, and combine the results. I went back and tested it. On my project dataset, it worked. I got lucky, very lucky: one short conversation produced an idea I could test right away, and the idea actually held up.</p>

<p>After those two projects, I understood why the first one had failed. As a software engineer, I was used to someone else having already defined the product, roadmap, or milestone. In a small research project, I was also the product manager. I had to choose the question, decide what evidence would count, and narrow the scope before the calendar narrowed it for me. Besides, research is inherently open-ended and much less deterministic than writing a software feature.</p>

<p>My internship and thesis forced the same issue at a larger scale. The work needed business value and an academic contribution, while the data available inside the company limited what I could honestly study. I gained much more respect for the need to define the right baselines and explicit guardrails, and to ask the annoying but necessary question: what evidence do I actually have?</p>

<p>Between classes, I could often be found at a CSAIL seminar or thesis defense that I had learned about through one of the many mailing lists I joined. I later wrote about <a href="/posts/2025/10/mit-cs-ai-engagement/">finding my place in MIT’s CS and AI community</a> and <a href="/posts/2025/12/ai-research-inside-nike/">conducting AI research inside Nike</a>.</p>

<p>I had entered MIT thinking mainly about technical depth: read more, learn more, build something difficult. All of that mattered. But I had underappreciated how unpredictable research could be. Today, I believe I have gotten better at recognizing when I am drifting, and sometimes I can catch it early enough to change direction.</p>

<figure>
  <a href="/images/2026-07-29-MIT-Reflection/Posters.jpeg">
    <img src="/images/2026-07-29-MIT-Reflection/Posters.jpeg" alt="A composite of Junru beside a poster on guardrailed AI for procurement negotiation and with two classmates beside a poster on social sycophancy in multi-turn AITA advice" width="1280" height="720" loading="lazy" />
  </a>
  <figcaption>Two research poster sessions.</figcaption>
</figure>

<h3 id="600-pm---party-or-pset">6:00 P.M. - Party or Pset?</h3>

<p>At the end of a class day, I often had two versions of the evening available.</p>

<p>One version continued with dinner, friends, and a Sloan C-Function (styled as <em><code class="language-plaintext highlighter-rouge">C-f(x)</code></em> when we wanted to look nerdy). I think the C stands for <a href="https://mitsloan.mit.edu/student-life/community-events#:~:text=synonymous%20with%20community">community</a>. Every other Thursday or so, student organizations put together a themed party, usually with food, drinks, and some sort of cultural program.</p>

<p>I played Chinese oldies on the clarinet at an Asian American Alliance C-Function. A wine club event brought in dozens of brands for a whole night of tasting. And just two weeks before final exams, I got to join more than 400 other students at Paradise Rock Club and cheer on our talented classmates as they rocked and rolled onstage at “Rolling Sloan.” Sometimes, simply heading to Middlesex, a local Cambridge establishment, and dancing for an hour with a small group of friends also counted as “living life to the fullest.”</p>

<p>If I went to a C-Function, I usually enjoyed it fully and went home to sleep. If I did not go, there was probably an engineering pset or exam waiting for me. That version of the evening generated far fewer social media posts.</p>

<p>I tried to experience the MBA social world, but I never came close to doing all of it. EECS work consumed time that some classmates used for clubs or broader MBA social life. But hey, I remembered what had brought me to grad school in the first place and made the necessary tradeoffs.</p>

<h3 id="1030-pm---late-night-mit">10:30 P.M. - Late-Night MIT</h3>

<p>I spent many late nights working on engineering assignments with fellow LGO students, non-LGO graduate students, and even undergraduates. These late-night sessions helped me get unstuck on a difficult concept and, more importantly, build new bonds with students outside my usual circles at MIT.</p>

<p>Then there was the Monday night when I lost my AirPods. The search produced perhaps the most MIT solution possible to a missing pair of earbuds.</p>

<p>After a weekend trip, I opened Find My and saw that they were somewhere inside MIT.nano. This was not especially helpful. MIT.nano has multiple floors and a clean room, while the dot on my phone could not identify an actual room. I started searching floor by floor.</p>

<p>My phone occasionally detected a weak signal, but it was not strong enough for precision finding, and I could not hear the AirPods playing a sound. Two undergraduates studying in a quiet area stopped what they were doing and helped me look.</p>

<p>An MIT.nano PhD student then learned what had happened. He had access to the clean room, which I could not enter, so he took my phone, went through the gowning process, and searched inside the clean room for about ten minutes. He came back without the AirPods, sadly.</p>

<p>The signal eventually became strongest near a locked office beside the clean-room entrance. A staff member had found the AirPods and put them in that office, which explained why my phone kept directing us toward the clean room. I left a handwritten note on the locked door and returned the next day to pick them up.</p>

<p>This was a ridiculous amount of MIT infrastructure and human effort for one pair of earbuds. I will always remember this funny, wholesome night.</p>

<p>All was quiet that night. I hopped back on my red bike and pedaled westward home.</p>

<p>And that wraps up a “typical day” for me.</p>

<figure style="max-width: 36rem; margin-left: auto; margin-right: auto;">
  <a href="/images/2026-07-29-MIT-Reflection/Bananas.jpeg">
    <img src="/images/2026-07-29-MIT-Reflection/Bananas.jpeg" alt="Junru in MIT graduation regalia holding bananas and a red diploma folder in front of stacked banana boxes" width="853" height="1280" />
  </a>
  <figcaption>One last snack from the Banana Lounge.</figcaption>
</figure>

<h2 id="trips-outside-my-usual-world">Trips Outside My Usual World</h2>

<p>A recurring theme of my MBA experience was exploring things I otherwise would not have encountered. Beyond classes, student-led trips produced some of my most memorable and shareable experiences. In addition to the plant treks organized by LGO, I joined two especially interesting trips.</p>

<p>On the <a href="https://www.linkedin.com/company/harvard-mit-germany-trek">Harvard-MIT Germany Trek</a> in August 2025, I had the good fortune to meet leaders in government, industry, and academia. In Erfurt, <a href="https://en.wikipedia.org/wiki/Mario_Voigt">Mario Voigt</a>, the minister-president of Thuringia<sup id="fnref:5" role="doc-noteref"><a href="#fn:5" class="footnote" rel="footnote">5</a></sup>, personally grilled sausages for us while discussing the region’s innovation and history. In Berlin, I greeted <a href="https://en.wikipedia.org/wiki/Christian_Wulff">Christian Wulff</a>, former president of Germany<sup id="fnref:6" role="doc-noteref"><a href="#fn:6" class="footnote" rel="footnote">6</a></sup>, in my beginner-level German and discussed international relations with him. Inside Germany’s Federal Chancellery<sup id="fnref:7" role="doc-noteref"><a href="#fn:7" class="footnote" rel="footnote">7</a></sup>, I asked the former head of Deutsche Bahn<sup id="fnref:8" role="doc-noteref"><a href="#fn:8" class="footnote" rel="footnote">8</a></sup> a question near and dear to me because of my love of trains: How could DB revive the operational excellence for which it was once renowned worldwide?</p>

<p>In March 2026, I joined an MBA Study Tour focused on heritage luxury-goods brands. In Milan, we were guided into Dolce&amp;Gabbana’s couture atelier, a palatial, museum-like space unmarked on maps, where its finest works and jewelry were displayed. In Sant’Agata Bolognese, I walked along Lamborghini’s assembly line and saw how the brand prioritized craftsmanship over volume throughput. Then, in Paris, Van Cleef &amp; Arpels took us into its VIC-only gallery behind Place Vendôme, which demonstrated step by step how a piece of high jewelry is meticulously made. Early one morning, we entered Hermès’s flagship maison on Faubourg Saint-Honoré before it opened to the public and heard the store’s head take us through the brand’s history and future. Finally, MIT Sloan alumnus Jean Arnault hosted us at LVMH, where I asked whether he would make a “smart LV watch.” Jean explained his view through an analogy: high watchmaking is like F1; he wanted all the technology and AI around the product, but not necessarily inside it.</p>

<p>Apparently, even on a luxury study tour, “Junru the Engineer” could not stop asking operations and product questions. I came back with a wider sense of what I could be curious about. That was one of the privileges of being a student again.</p>

<!--
## A Leadership Reflection

Leadership was a major component of my MIT program, as it is in most business school programs. It is an abstract word with many interpretations and associated practices. Business school students are expected to already have some impressive "leadership moments" because every application asks for them, but what does leadership really mean? I have pondered this a lot since I started the program.

For a long time, I understood leadership mainly through connection. I wanted to build relationships across languages, cultures, industries, and functions. I still recognize myself in that definition, but it now feels incomplete.

During an international plant trek in China, I informally guided more than ten classmates through a cherry-blossom park and helped with taxis and local travel logistics. It was a small thing. For a few hours, though, I could help people feel comfortable enough to be curious and to explore in a space largely foreign to them. That day felt close to the kind of leadership I believed in.

Another program trip went very differently. A few classmates and I had inherited responsibility for an anonymous social-media account that documented a cohort tradition. Many classmates enjoyed the photos and jokes that the other "secret historians" and I curated. I was already known as someone who photographed and documented our shared experiences, so the role felt natural to me.

Other classmates were uncomfortable because they did not know who controlled the account. LGO alumni from previous years followed it, and posts could travel into professional or recruiting circles. The same photo could feel like a community memory to one person and unwanted exposure to another.

By the end of the trip, frustration within the class had escalated and was directed at all the "secret historians" behind the account. I felt as though I had been put on public trial.

Those were the most painful few days of my two years at MIT.

I stepped forward and took responsibility for the account.

My first reaction was that I could never please everyone. That was true, but it also let me off too easily. I had mistaken some classmates' appreciation for consent from the whole group. A better setup would have made clear who was responsible, set boundaries in advance, and given everyone a way to opt out.

I still want to connect people. But I have become more mindful of boundaries and the different ways in which people operate. Leadership also requires me to think about consent, take responsibility for the boundaries around a shared space, and notice what my attempt at connection might cost someone else. I understood this only after the conflict became painful.

I was also lucky to have many gracious classmates and friends who supported me as I navigated this mini-crisis. Years from now, this "crisis" will probably feel like "nothing," but it was nonetheless a moment that shaped my approach to leadership for the better.
-->

<h2 id="was-it-worth-it">Was It Worth It?</h2>

<p>Would I choose LGO again?</p>

<p>Yes, with conditions.</p>

<p>It was worth it for me because I wanted technical depth and also knew that I did not want to remain solely a software engineer for the rest of my career. I was curious about products, operations, organizations, leadership, and the workings of many other industries, even though I did not have a precise title for the person I hoped to become.</p>

<p>The combination had tradeoffs. Every difficult engineering project took time away from the broader MBA social world. LGO made the dual degree navigable, but I still had to find a research community and construct my own technical identity. The program supplied the skeleton and left me to decide how to use the space inside it.</p>

<p>I still think someone who wants only technical depth would be better served by a focused master’s or PhD path. LGO is often reduced to an “MS + MBA,” but that shorthand leaves out what makes the program distinct: its focus on operations and its unique portfolio of partner companies that fund the program to develop future operations talent.</p>

<p>As explained earlier, I did not apply because I dreamed of a traditional operations career. I expected to return to tech. Even before MIT, though, I had started noticing that a powerful algorithm was not enough to make a product work well in the real world. LGO gave me more ways to think about everything around the algorithm, especially what must happen before a technical idea becomes a reliable product. I do not think every applicant has to want a job managing a factory. But the factory, the supply chain, and all the other equally important work around a product cannot feel like someone else’s problem.</p>

<p>For my own pre-MIT self, the answer is clear: I would choose LGO again.</p>

<figure>
  <a href="/images/2026-07-29-MIT-Reflection/LGO_Lounge.jpeg">
    <img src="/images/2026-07-29-MIT-Reflection/LGO_Lounge.jpeg" alt="Junru in a suit and LGO graduation stole relaxing with his feet on a desk in the LGO student lounge" width="1280" height="853" loading="lazy" />
  </a>
  <figcaption>Me in 2026, endorsing my past decisions :) Taken in the LGO student lounge, where we spent days and nights working, drinking free coffee, and emptying each month's snack stash only a few days after it arrived.</figcaption>
</figure>

<p>The program did not send me in a completely new direction. I arrived interested in technically difficult products, and I still am. I now ask earlier what problem the product actually solves, what evidence I have, and how a decision lands on the people affected by it. I am also more willing to admit when the question itself is still unclear.</p>

<p>I came to MIT mainly because I wanted a technical reset. I left with a stronger technical foundation, more practice deciding what to ignore, and more caution about decisions that affect other people.</p>

<p>Years from now, if this post helps me picture the French bike on Vassar Street or remember waiting outside that locked MIT.nano office, it will have done its job.</p>

<p>Lastly, thank you to everyone who showed me grace and kindness: my family and loved ones; friends and classmates across LGO, Sloan, and EECS; professors, TAs, advisors, coaches, and staff; Nike teammates, former colleagues at SoundHound, and LGO alumni; and the many strangers who helped along the way. And thank you to my red French bike, my second-hand sedan, and the trusty MBTA for literally carrying me through much of it.</p>

<figure style="max-width: 42rem; margin-left: auto; margin-right: auto;">
  <a href="/images/2026-07-29-MIT-Reflection/LGO26_First_and_Last_Day_in_School.jpeg">
    <img src="/images/2026-07-29-MIT-Reflection/LGO26_First_and_Last_Day_in_School.jpeg" alt="Two group portraits of the LGO Class of 2026, at the beginning of the program and wearing graduation stoles at its end" width="960" height="1280" loading="lazy" />
  </a>
  <figcaption>The LGO Class of 2026, from our first day at MIT on June 3, 2024, to our last on May 28, 2026.</figcaption>
</figure>

<h2 id="appendix-two-years-by-the-numbers">Appendix: Two Years by the Numbers</h2>

<ul>
  <li>📚 Completed 231 graduate units apart from the thesis</li>
  <li>☕️ Drank nearly 2,000 cups of coffee</li>
  <li>🏭 Visited 11 companies across 7 states during one 21-day domestic plant trek</li>
  <li>🚚 Moved cross-country 4 times: CA to MA, MA to OR, OR back to MA, and finally MA to CA</li>
  <li>✈️ Traveled to the following destinations (excluding short layovers):
    <ul>
      <li>18 U.S. states: Massachusetts, Oregon, Washington, Maine, New Hampshire, Vermont, Rhode Island, New Jersey, New York, Pennsylvania, Georgia, Florida, Tennessee, Louisiana, Illinois, Utah, Arizona, and California</li>
      <li>5 countries outside the U.S.: China, Thailand, Germany, France, and Italy</li>
    </ul>
  </li>
  <li>🗣️ Started learning 3 languages on Duolingo (German, French, and Italian), though I remain very much a beginner in all three</li>
  <li>💒 Officiated or emceed 2 weddings</li>
  <li>👟 Ran 388.25 miles (not an awful lot) but finally got into running thanks to Nike and the many hardcore runners in my class</li>
  <li>🚲 Biked the Minuteman Bikeway only twice</li>
  <li>🏟️ Went to 2 Red Sox games, 1 Celtics game, and 1 Bruins game, but never made it to a Patriots or Revolution game</li>
  <li>🚉 Rode all kinds of MBTA rail transit lines except the Blue Line</li>
  <li>⛵ Held a valid MIT Sailing Card for 2 seasons and sailed 12 times, yet never earned my <a href="https://sailing.mit.edu/card/provisional.php">Provisional Rating</a></li>
</ul>

<figure style="max-width: 36rem; margin-left: auto; margin-right: auto;">
  <a href="/images/2026-07-29-MIT-Reflection/Sailing.jpeg">
    <img src="/images/2026-07-29-MIT-Reflection/Sailing.jpeg" alt="Junru in a suit and LGO graduation stole standing beside a docked MIT dinghy on the Charles River" width="853" height="1280" loading="lazy" />
  </a>
  <figcaption>Two sailing seasons, and still no provisional rating.</figcaption>
</figure>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:1" role="doc-endnote">
      <p>They were (1) <a href="https://vimeo.com/1185687526?share=copy&amp;fl=cl&amp;fe=ci#t=4664.747">School of Engineering &amp; Schwarzman College of Computing Advanced Degree Ceremony</a>, (2) <a href="https://www.youtube.com/watch?v=Swdt8v4O1Mo">OneMIT Commencement Ceremony</a>, (3) LGO Convocation, and (4) Sloan School of Management MBA and Master of Science Management Studies Degree Ceremony. <a href="#fnref:1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:2" role="doc-endnote">
      <p>Source: <a href="https://news.mit.edu/2012/mit-lgo-logs-25-years">https://news.mit.edu/2012/mit-lgo-logs-25-years</a>. <a href="#fnref:2" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:3" role="doc-endnote">
      <p>See LGO’s <a href="https://lgo.mit.edu/academics/">program structure and timeline</a>. <a href="#fnref:3" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:4" role="doc-endnote">
      <p>What’s a <a href="https://mitsloan.mit.edu/mba/explore-program/mba-curriculum#:~:text=Following%20MIT%20Sloan%20tradition%2C%20each%20cohort%20is%20named%20after%20a%20body%20of%20water%3A%20Atlantic%2C%20Baltic%2C%20Caribbean%2C%20Indian%2C%20Mediterranean%2C%20and%20Pacific">cohort/ocean</a>? <a href="#fnref:4" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:5" role="doc-endnote">
      <p><em>Ministerpräsident des Freistaats Thüringen</em>: Minister-President of the Free State of Thuringia, roughly analogous to a U.S. state governor. <a href="#fnref:5" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:6" role="doc-endnote">
      <p><em>Bundespräsident</em>: Federal President, Germany’s head of state, not to be confused with the Federal Chancellor, who is the head of government. <a href="#fnref:6" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:7" role="doc-endnote">
      <p><em>Bundeskanzleramt</em>: Federal Chancellery, the institution that supports the Federal Chancellor and coordinates federal government policy; the term also refers to its Berlin building. <a href="#fnref:7" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:8" role="doc-endnote">
      <p>The national railway company of Germany. <a href="#fnref:8" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><category term="personal" /><category term="mit" /><category term="mit-sloan" /><category term="grad-school" /><summary type="html"><![CDATA[Two months ago, in May 2026, I somehow graduated from MIT four times.1 Somewhere in that run of ceremonies, in front of the Great Dome overlooking the Charles River, I listened to Dr. Lisa Su’s inspiring speech about being ambitious, solving hard problems, and finding ways to make luck. I also snacked one last time on bananas from the Banana Lounge. The final result was two degrees, an S.M. in EECS and an MBA, through the Leaders for Global Operations (LGO) program with a $100,000 fellowship. They were (1) School of Engineering &amp; Schwarzman College of Computing Advanced Degree Ceremony, (2) OneMIT Commencement Ceremony, (3) LGO Convocation, and (4) Sloan School of Management MBA and Master of Science Management Studies Degree Ceremony. &#8617;]]></summary></entry><entry><title type="html">Jazz and Homesickness</title><link href="https://junruren.com/posts/2025/12/jazz-and-homesickness/" rel="alternate" type="text/html" title="Jazz and Homesickness" /><published>2025-12-31T00:00:00+00:00</published><updated>2025-12-31T00:00:00+00:00</updated><id>https://junruren.com/posts/2025/12/jazz-and-homesickness</id><content type="html" xml:base="https://junruren.com/posts/2025/12/jazz-and-homesickness/"><![CDATA[<p>On a cold January day in 2023 in Nashville, Tennessee, just before heading to the airport, I took a small detour to stop by Nashville Jazz Workshop. The venue was closed, but I stood outside for a moment, feeling as if a smooth piano melody were drifting through the air, quietly comforting my homesick heart. If this pairing of American jazz and my Chinese homesickness catches your attention, I invite you into a very personal story of mine, one that intertwines the life journey of an international student with a particular American jazz pianist. Happy 2026, from my home in Shenzhen, China.</p>

<p class="notice">This post also lives on <a href="https://junruren.substack.com/p/jazz-and-homesickness?utm_source=junruren.com&amp;utm_medium=referral&amp;utm_campaign=jazz-and-homesickness">my Substack</a> — comment there, or subscribe to get future posts by email.</p>

<hr />

<h2 id="first-tastes-of-homesickness">First Tastes of Homesickness</h2>

<p>How I decided to study abroad in the United States is a story in itself and probably warrants a separate post.</p>

<p>My study-abroad journey unfolded quite smoothly in sunny San Diego. I quickly made friends, adapted to the local lifestyle, and became active in various community activities. Outside of school, I was blessed with <a href="https://convoydistrict.com">San Diego’s easy access</a> to groceries and cuisines from back home whenever I craved them. My family was also just one video call away, and we caught up almost every weekend.</p>

<p>For the first time in my life, I felt truly independent and empowered. Life was great in San Diego.</p>

<p>I did not really feel homesick, until two things happened.</p>

<p>At the end of my freshman year, everyone living on campus had to vacate their dorm by the Sunday following finals week. I had sublet a place for the summer, and my only task was to move my belongings from my dorm room to the new apartment, a short ten-minute drive away. While I was licensed to drive, I did not own a car at the time, and being underage meant that I could only use whatever Zipcar offered on campus. I reserved a Toyota Prius for the day.</p>

<p><img src="/images/2025-12-31-jazz-and-homesickness/Inside-Move-Out-Prius.JPG" alt="Inside the Prius" /></p>
<blockquote>
  <p>Inside the Prius I rented. I am surpised that I actually took a picture back then.</p>
</blockquote>

<p>On move-out day, the residential area felt more crowded than when school was in session. Family members of local students showed up in force: parents, grandparents, and siblings arrived in minivans or pickup trucks, equipped with flatbed wagons and sturdy Home Depot boxes, ready to get the job done quickly. I had a lovely interaction with my roommates’ family, and before I knew it, their move-out was complete.</p>

<p><img src="/images/2025-12-31-jazz-and-homesickness/Dorm-Vacated.jpg" alt="Dorm room vacated" /></p>
<blockquote>
  <p>My vacated dorm room, snapped on Snapchat (ah the Snapchat years…)</p>
</blockquote>

<p>Meanwhile, I had just bought a folding hand truck dolly cart. Together with the Prius, it meant making multiple trips between my dorm and the new apartment. While driving off campus, I spotted another international-student friend who was about to move their luggage by taking a local bus. I offered a ride.</p>

<p>By sundown, I had paused my own move and given ride after ride to friends who needed help. My Prius was used to its fullest capacity. Looking back now, I am fairly certain that the nineteen-year-old version of me could have packed and managed time more efficiently. My hauling was not finished until well past midnight. I had not eaten dinner, so I drove to a Subway and took my first bite of the night. It was at that moment, with a foot-long sub in hand at 1 a.m., that I felt the “curse” of independence: I did not have my parents in a minivan. I was on my own.</p>

<p>The other moment of missing home also came toward the end of my freshman year. I suffered a severe bout of food poisoning and was eventually taken by ambulance to the emergency room. After spending the night in the hospital by myself, I grabbed an Uber and went straight to my first lecture of the morning.</p>

<p><img src="/images/2025-12-31-jazz-and-homesickness/ER.JPG" alt="Inside UC San Diego Health ER" /></p>
<blockquote>
  <p>Inside UC San Diego Health ER, my first taste of American healthcare.</p>
</blockquote>

<blockquote>
  <p><em>Sidebar</em>: I recently listened to <a href="https://fortune.com/2025/12/05/jensen-huang-nine-year-old-janitor-before-nvidia-joe-rogan-podcast/">Jensen Huang’s story</a>, in which he described communicating with his parents through a tape cassette. The story is deeply touching, and it reminded me not to take the support I have for granted.</p>
</blockquote>

<p>I learned that vulnerable moments make us cherish the love we had previously taken for granted. Although I was initially proud of how easily I had adapted to life abroad, that sense of triumph faded as soon as the first real setback hit.</p>

<p>After my freshman year, whenever I flew home to reunite with my family, <strong>I began searching for something random yet tangible that could remind me of home</strong>.</p>

<h2 id="trans-pacific-commutes">Trans-Pacific Commutes</h2>

<p>Between my hometown of Shenzhen and my college town of San Diego, flights operated by Cathay Pacific between Hong Kong and Los Angeles became my go-to route for holidays and returns to school. After those early homesick moments, I started associating a Cathay Pacific flight with going back home (regardless of the direction of a flight).</p>

<p>On board a home-bound flight in my sophomore year, smooth jazz piano flowed through the in-flight entertainment system into my headset Curious, I tapped through the IFE screen in front of me. The artist was Beegie Adair.</p>

<p>Elegant, melodic, sophisticated, yet approachable. That was Beegie’s musical style.</p>

<p>The album playing was <a href="https://beegieadair.com/album/338049/cocktail-party-jazz-2">“Beegie Adair &amp; Friends: Cocktail Party Jazz 2”</a>. Some might label it “lounge music” (and it certainly can be), but that first immersion carried me through a quietly sentimental journey <sup id="fnref:1" role="doc-noteref"><a href="#fn:1" class="footnote" rel="footnote">1</a></sup>.</p>

<p><img src="/images/2025-12-31-jazz-and-homesickness/IFE-with-Beegie.jpeg" alt="Actual photo I took of the IFE screen with the exact album" /></p>
<blockquote>
  <p>Actual photo I took of the IFE screen with the exact album. I cannot believe I actually took this photo on that flight. (Well, I sure took it so that I could search for the album after landing.)</p>
</blockquote>

<p>By the time the wheels touched down, the same album was still rolling on my IFE.</p>

<p><strong>Random but tangible</strong>: from that point on, hearing Beegie Adair’s music became associated, in my mind, with being on board a Cathay flight, and being on board a Cathay flight became associated with going home. It is difficult to explain how my mind formed this two-stage connection, though I likely reinforced it by deliberately replaying the same album on subsequent Cathay flights, regardless of the direction of travel. The impression from that first listen was simply too strong.</p>

<h2 id="into-jazz">Into Jazz</h2>

<p>Jazz, in its broadest sense, was part of my childhood soundscape. My dad often played Kenny G in the car, which nudged me toward learning a woodwind instrument (clarinet, not Kenny’s sax, somehow), and I was trained classically. I never actively sought out more jazz, but I welcomed it whenever I encountered it.</p>

<p>That changed after the inexplicable establishment of my Beegie’s jazz == homesickness cure logic. Through her recordings, I began to explore jazz standards and the Great American Songbook. I had entered a new musical world.</p>

<p>After college, I was fortunate to join a company full of musicians who loved jazz. We jammed in the office music room every Tuesday and Thursday after work. It remains my best work memory.</p>

<p>One quiet wish stayed with me: to hear Beegie perform live one day. I knew she played regularly at <a href="https://www.nashvillejazz.org">Nashville Jazz Workshop</a>, where she was also a founding board member. When my post-graduate life finally allowed for more travel after a year of work, the COVID-19 pandemic arrived, pushing my Nashville plans indefinitely into the future.</p>

<p>Through the thick of lockdowns and my own medical concerns, I eventually reached a point where I felt ready to look up Beegie’s next performance. Instead, I found a solemn line on <a href="https://www.beegieadair.com">https://www.beegieadair.com</a>:</p>

<blockquote>
  <p>January 23, 2022: It is with deep and profound sadness we share the sad news that Beegie died today surrounded by those she loved and who loved her most dearly…</p>
</blockquote>

<h2 id="a-tribute">A Tribute</h2>

<p>My first visit to Nashville finally came in January 2023. That trip had more nuance that I may touch on in a future reflection, but amid it, I was in Nashville with a sense of regret that I could never close my “jazz-homesick loop” by hearing Beegie play live anymore.</p>

<p>On my Uber ride to the airport, I added Nashville Jazz Workshop as a stop. I stood outside the building for a few seconds, took a selfie, and then returned to my airport-bound ride. I put on my headphones and pressed play on <em>Cocktail Party Jazz 2</em>.</p>

<p><img src="/images/2025-12-31-jazz-and-homesickness/Nashville-Jazz-Workshop.jpeg" alt="Selfie taken outside Nashville Jazz Workshop" /></p>
<blockquote>
  <p>Selfie taken outside Nashville Jazz Workshop</p>
</blockquote>

<hr />

<p>Today, we each consume information on the order of gigabytes per day. Countless random details pass through our lives. Yet somehow, just somehow, I formed an unexplainable set of connections among a few seemingly unrelated random things: a musician, an airline, and a feeling of homesickness. Over time, that linkage quietly amplified its own significance in my mind.</p>

<p>I appreciate you reading this far because you probably bore the desire to shout: how on earth did I come to make these connections… Well, I would ask that question of myself too. But here I am: I’ve invented and kept within me a unique embodiment of some very layered emotions and memories, beyond the economy cabin of a trans-Pacific flight carrying an anxious international student away from home.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:1" role="doc-endnote">
      <p>A small pun intended: “Sentimental Journey” is another jazz standard that Beegie also performed. <a href="#fnref:1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><category term="personal" /><category term="music" /><summary type="html"><![CDATA[On a cold January day in 2023 in Nashville, Tennessee, just before heading to the airport, I took a small detour to stop by Nashville Jazz Workshop. The venue was closed, but I stood outside for a moment, feeling as if a smooth piano melody were drifting through the air, quietly comforting my homesick heart. If this pairing of American jazz and my Chinese homesickness catches your attention, I invite you into a very personal story of mine, one that intertwines the life journey of an international student with a particular American jazz pianist. Happy 2026, from my home in Shenzhen, China.]]></summary></entry><entry><title type="html">What is it like to do AI research inside Nike?</title><link href="https://junruren.com/posts/2025/12/ai-research-inside-nike/" rel="alternate" type="text/html" title="What is it like to do AI research inside Nike?" /><published>2025-12-20T00:00:00+00:00</published><updated>2025-12-20T00:00:00+00:00</updated><id>https://junruren.com/posts/2025/12/AI-Research-inside-Nike</id><content type="html" xml:base="https://junruren.com/posts/2025/12/ai-research-inside-nike/"><![CDATA[<p>I’m feeling bittersweet because yesterday, December 19, 2025, was my last day at Nike. My six-month research internship culminated in three presentations across multiple time zones. While the details of the work will eventually show up in my MIT thesis next year, here I reflect on this unique journey of conducting AI research at a world-renowned brand.</p>

<p class="notice">This post also lives on <a href="https://junruren.substack.com/p/ai-research-inside-nike?utm_source=junruren.com&amp;utm_medium=referral&amp;utm_campaign=ai-research-inside-nike">my Substack</a> — comment there, or subscribe to get future posts by email.</p>

<p><img src="/images/2025-12-20-AI-Research-inside-Nike/Nike-Campus-Last-Day.jpeg" alt="Amazing sunset over Lake Nike" /></p>
<blockquote>
  <p>On my last day, the rainy Oregon sky cleared up just in time for a stunning sunset over Nike’s <a href="https://about.nike.com/en/newsroom/releases/nike-philip-h-knight-campus-announcement">Philip H. Knight Campus</a></p>
</blockquote>

<h2 id="the-dual-mandate">The Dual Mandate</h2>

<p><strong>You must think about complicated business needs and state-of-the-art research at the same time.</strong> At a company like Nike, the first priority is being intentional about how the work solves real problems and creates wins. For example, if I get too enamored with a valid problem in academia that’s not yet relevant to the business, it might not create immediate value. This is exactly why I’m grateful to bring both my MIT EECS and MBA lenses to the job.</p>

<p><strong>Networking is a must.</strong> In any huge company, knowledge of a complex business process is naturally distributed. The most effective way (in my experience) to connect the dots is to form collaborative partnerships with teammates who have the subject-matter expertise. Map out everything, put it into one complete document, invite feedback, and iterate.</p>

<p><strong>Literature review starts on day one.</strong> When your research takes the form of an industry internship, it’s easy to lose sight of what’s happening in academia, especially in the midst of all the networking above. I’m grateful my thesis advisors constantly reminded me to find the papers that back up what I’m seeing (or claiming). There’s no real substitute for shuttling between “professional work” and “research.” This context switching happens basically every other day.</p>

<h2 id="ai-at-nike">AI at Nike</h2>

<p>From the outside, Nike is not likely associated with AI research. However, the company is deeply invested in leveraging AI to enhance customer experiences, optimize supply chains, and innovate product designs. During my internship, I witnessed firsthand how AI is integrated into various facets of Nike’s operations. One recent example that the public may be familiar with is the <a href="https://www.thestack.technology/nikes-cto-welcomes-nikeai-rollout-are-there-lessons-here-for-developers/">AI search feature on the Nike app</a>, which allows users to find products in natural language.</p>

<p><strong>My first taste of an enterprise AI platform.</strong> Nike is the very first large enterprise that I have ever worked at (naturally given that my only previous full-time endeavor was at an AI startup), and I truly didn’t know much about how data is managed at scale and whether there is an established platform on which AI solutions can be developed.</p>

<p>Through this internship, I got my hands on Databricks for the first time. On the surface, Databricks is a data lakehouse platform that unifies data engineering, data science, and business analytics. But quickly I realized that it is also a powerful AI platform that allows researchers and engineers to build, deploy, and monitor machine learning models at scale. From a simple regression model to access to nearly all major large language models (LLMs) via API, Databricks provides a seamless environment for AI research and development within a large organization like Nike.</p>

<p>Need to embed some text data? Databricks has you covered with various embedding models. Done training your model? You can easily deploy it on Databricks and create an endpoint for real-time predictions. My amazement grew as I saw how Databricks quickly made a new model like GPT-5.2 available as soon as OpenAI released it. And more importantly, all of these models are accessed in a secure and compliant manner, which is crucial for enterprise applications.</p>

<p>When I was just a student/individual builder outside of big companies, it rarely occurred to me how important internal data governance is; sharing my whole life’s troubles with the public instance of ChatGPT.com seems like the way to go. (Well, I shouldn’t…)</p>

<p><strong>I was also blessed with a supportive environment.</strong> From my day-to-day data science teammates to the VP of the customer organization that my research serves, everyone was incredibly supportive and encouraging. They provided me with the resources, mentorship, and feedback needed to thrive in this internship. Over the past six months, I interacted with a far more diverse set of functions than I would have at a purely technical company. In AI, my technical mentors are all hands-on builders who ship models or agentic tools every day, and my business function partners are genuinely excited about applying AI to solve their real problems.</p>

<h2 id="cool-company">Cool Company</h2>

<p>If you know Nike, you know it’s a cool company. The brand is iconic, the culture is vibrant, and the people are passionate about what they do—and about sports. I warmed up with <a href="https://www.facebook.com/share/v/1H299taBo3/">Eliud Kipchoge before a 5K on campus</a>, and I’ve made sports a daily habit by frequenting three different on-site gyms. Running, cycling, swimming, repeat. I grew up with the Swoosh too, so it’s kind of magical to work behind a brand that’s instantly recognized around the world. I also joined during an important turning point as the company sharpens its focus on winning again, winning now, so it was valuable to witness a major transition unfold—and to watch how people from senior leadership to every individual contributor react and adapt.</p>

<p>If you ask me what the coolest experience during my internship was, I’d say it was the <strong>Shoe School</strong>, where I got to hand-make my own shoe from scratch! Sadly, it was just one for the left foot, because the teaching materials were ordered in batches such that each cohort gets to make either a left shoe or a right shoe. Still, it was an unforgettable experience to understand the craftsmanship and technology that goes into making a Nike shoe.</p>

<p><img src="/images/2025-12-20-AI-Research-inside-Nike/Shoe-School.jpeg" alt="I made a shoe!" /></p>

<p>At Nike, there is also a <strong>unique MIT community.</strong>  My research internship at Nike was part of the MIT LGO program, and this <a href="https://lgo.mit.edu/partnercompanies/nike-inc/">~15-year partnership</a> has created a strong alumni network on campus. I was lucky to have this community’s support: both for career development and for fun activities.</p>

<p>And yes, I’ll admit it: I’ve collected a dangerous number of Nike shoes (free or at huge discounts, thanks to the employee perks). But hey, I’m running comfortably in my brand new Vomero Plus, and my knees will be happier.</p>

<p><img src="/images/2025-12-20-AI-Research-inside-Nike/Shoe-Boxes.jpeg" alt="My Nike Shoe Boxes" /></p>

<hr />

<p>Well, finally, a teaser on my research: <strong>a multi-agent buyer–seller negotiation framework under limited information</strong>…</p>

<p>Enough said. <a href="/posts/2025/08/MIT-Thesis-LaTeX/">LaTeX time</a>…</p>]]></content><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><category term="research" /><category term="ai" /><category term="internship" /><category term="nike" /><category term="mit" /><summary type="html"><![CDATA[I’m feeling bittersweet because yesterday, December 19, 2025, was my last day at Nike. My six-month research internship culminated in three presentations across multiple time zones. While the details of the work will eventually show up in my MIT thesis next year, here I reflect on this unique journey of conducting AI research at a world-renowned brand.]]></summary></entry><entry><title type="html">Stolen Name: How Software Erases Identity</title><link href="https://junruren.com/posts/2025/11/stolen-name/" rel="alternate" type="text/html" title="Stolen Name: How Software Erases Identity" /><published>2025-11-04T00:00:00+00:00</published><updated>2025-11-04T00:00:00+00:00</updated><id>https://junruren.com/posts/2025/11/Stolen-Name</id><content type="html" xml:base="https://junruren.com/posts/2025/11/stolen-name/"><![CDATA[<p>Names are one of the first things a system asks for, and one of the easiest ways it reveals what culture it was built for. This post looks at how seemingly harmless assumptions in software can quietly erase identity, using Chinese names as the primary case study.</p>

<p class="notice">This post also lives on <a href="https://junruren.substack.com/p/stolen-name?utm_source=junruren.com&amp;utm_medium=referral&amp;utm_campaign=stolen-name">my Substack</a> — comment there, or subscribe to get future posts by email.</p>

<p><img src="/images/2025-11-04-Stolen-Name/Cover_by_ChatGPT.png" alt="Cover image by ChatGPT" /></p>

<hr />

<p>In my junior years as a software engineer designing APIs to relay user information for voice AI, my mentor, <a href="https://www.ellipsix.net/index.html">David</a>, shared a classic post: <a href="https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-believe-about-names/"><em>Falsehoods Programmers Believe About Names</em></a>. It was humorously eye-opening at the time. My favorite is the last one:</p>

<blockquote>
  <p><em>40. People have names.</em></p>
</blockquote>

<p>We often model “name” as a tidy schema, then watch reality break it.</p>

<p>As someone who has built user-facing systems, I notice how the design of something as small as a name field encodes an entire worldview. Details such as whitespace, commas, and capitalization become cultural choices with real human consequences.</p>

<h2 id="the-issue-when-given-names-have-spaces">The Issue: When Given Names Have Spaces</h2>

<p>Take my given name: <strong>Junru</strong>.</p>

<p>On my passport, my name is romanized as “Ren, Junru.” That works fine in most Western systems: two tokens, mapped cleanly to family and then given names. When a service greets me, I usually see “Hi Junru!” No drama.</p>

<p>In reality, my given name is formed by two characters, Jun and Ru, as is common for many Han Chinese names. Mainland documents romanize those two characters as one token (“Junru”). Outside mainland China, though, it is common to romanize with a space (<strong>“Jun Ru.”</strong>) or with a hyphen (<strong>“Jun-Ru”</strong>).</p>

<p>Here’s where the breakage starts. Many systems assume a first-space split and misinterpret the space as a boundary between first and middle names, illustrated in this simple Python snippet:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">first_name</span> <span class="o">=</span> <span class="n">full_name</span><span class="p">.</span><span class="n">split</span><span class="p">(</span><span class="s">" "</span><span class="p">)[</span><span class="mi">0</span><span class="p">]</span>
<span class="n">last_name</span>  <span class="o">=</span> <span class="n">full_name</span><span class="p">.</span><span class="n">split</span><span class="p">(</span><span class="s">" "</span><span class="p">)[</span><span class="o">-</span><span class="mi">1</span><span class="p">]</span>
</code></pre></div></div>

<p>Neat. Deterministic. And, for “Jun Ru Ren,” it yields <code class="language-plaintext highlighter-rouge">first_name = "Jun"</code> and <code class="language-plaintext highlighter-rouge">last_name = "Ren"</code>, silently truncating half of the given name. As a result, one might be greeted as “Hi Jun!” instead of “Hi Jun Ru!” When others look up this name in a directory, they might see “Jun Ren”. A small bug becomes a small erasure.</p>

<p>This happens to friends who grew up in North America with given names romanized as two words. The downstream effects are social as much as technical: <strong>some eventually go by the truncated first token</strong>; others adopt an English name to avoid constant friction. Technology didn’t force the choice, but it nudged it.</p>

<h2 id="beyond-chinese-names-a-global-pattern">Beyond Chinese Names: A Global Pattern</h2>

<p>The point isn’t that Chinese names are uniquely tricky; it’s that a single schema can’t represent the world:</p>

<ul>
  <li>Some Indonesians and Burmese people have mononyms: no family name at all.</li>
  <li>Icelandic names are primarily patronymic/matronymic, not stable family surnames.
    <ul>
      <li>A favorite example: the Icelandic-Chinese jazz superstar <a href="https://en.wikipedia.org/wiki/Laufey_(singer)">Laufey</a>. Her full name is <strong>Laufey Lín Bing Jónsdóttir</strong> where <strong>Jónsdóttir = Jón + s + dóttir</strong>, meaning “daughter of Jón.” Her Chinese name 林冰 (Lin, Bing) is proudly included, with Lin being her mother’s family name. Interestingly, her twin sister <strong>Júnía Lín Hua Jónsdóttir</strong> appears as “Júnía Lin” in production credits.</li>
    </ul>
  </li>
  <li>Many Spanish-speaking cultures use two family names (paternal and maternal).</li>
  <li>In parts of India, name order and presence of a family name vary widely by region and language.</li>
</ul>

<p>Yet many forms still require “First Name” and “Last Name,” full stop.</p>

<p>A helpful resource for developers trying to do better is the W3C write-up, <a href="https://www.w3.org/International/questions/qa-personal-names">Personal Names Around the World</a>. The short version: accept variability, avoid premature parsing, and don’t assume structure from whitespace.</p>

<h2 id="appendix-how-chinese-ids-treat-names-chinese-script">Appendix: How Chinese IDs Treat Names (Chinese Script)</h2>

<p>When names are written in Chinese characters, the segmentation problem mostly disappears. Across most local ID systems, the full name appears as a single string of characters without explicit “first/last” fields.</p>

<table>
  <thead>
    <tr>
      <th>Sample</th>
      <th>Remark</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><img src="https://upload.wikimedia.org/wikipedia/commons/e/e7/The_People%27s_Republic_of_China_resident_identity_card_%28SAMPLE%29.png" alt="Sample image of a Resident Identity card in Mainland China (Source: Wikimedia)" /></td>
      <td><strong>Mainland China</strong> Resident ID Card shows one “full name” (姓名) field with characters together. In the sample photo, look for “某某某” as the placeholder of a full name. Some ethnic minority formats differ; see <a href="https://en.wikipedia.org/wiki/Resident_Identity_Card#Additional_features_in_ethnic_minority_areas">Additional Features</a></td>
    </tr>
    <tr>
      <td><img src="https://upload.wikimedia.org/wikipedia/commons/f/f6/ROC_mibunsho.jpg" alt="Sample image of a ROC Resident Identity Card (Source: Wikimedia)" /></td>
      <td>ID in <strong>Taiwan</strong> likewise shows one name field (姓名). Spacing between characters is typographic, not a given/family separator.</td>
    </tr>
    <tr>
      <td><img src="https://upload.wikimedia.org/wikipedia/commons/4/47/Hong_Kong_ID_card_front_side.png" alt="Sample image of a Hong Kong Permanent Identity Card (Source: Wikimedia)" /></td>
      <td><strong>Hong Kong</strong>: the Chinese line shows characters together; the English line separates with a comma between family and given names.</td>
    </tr>
    <tr>
      <td><img src="https://upload.wikimedia.org/wikipedia/commons/b/bf/MacaoID2023.jpg" alt="Sample image of a Macau Permanent Identity card (Source: Wikimedia)" /></td>
      <td><strong>Macau</strong> is a notable case where the card’s design explicitly separates family and given name fields.</td>
    </tr>
  </tbody>
</table>

<p>These examples (Macau aside) illustrate a key idea: in the native script, names are treated as a single unit; the friction arises during romanization and in systems that overfit to another schema.</p>

<h2 id="flip-the-table">Flip the Table</h2>

<p>We’ve looked at what happens when Chinese names meet systems designed around Anglo‑American conventions. Now imagine the reverse: a non‑Chinese name forced into a Chinese schema with no space‑based segmentation, or an interface that insists on family‑name‑first without a clear family name to give. It could go both ways.</p>

<p>If you have stories, what worked, what broke, I’d love to learn from them.</p>

<hr />

<h2 id="closing-thoughts">Closing Thoughts</h2>

<p>Whether in code or culture, the smallest design decisions shape how people see themselves. A name is not just data to be parsed; it’s a story, often written across languages and generations. When we build software, we are choosing which stories appear whole. Next time you see a field labeled “First Name”, pause and ask: whose first name?</p>]]></content><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><category term="software" /><category term="culture" /><summary type="html"><![CDATA[Names are one of the first things a system asks for, and one of the easiest ways it reveals what culture it was built for. This post looks at how seemingly harmless assumptions in software can quietly erase identity, using Chinese names as the primary case study.]]></summary></entry><entry><title type="html">Networking Beyond Your Home Department: An MIT CS/AI Case Study</title><link href="https://junruren.com/posts/2025/10/mit-cs-ai-engagement/" rel="alternate" type="text/html" title="Networking Beyond Your Home Department: An MIT CS/AI Case Study" /><published>2025-10-14T00:00:00+00:00</published><updated>2025-10-14T00:00:00+00:00</updated><id>https://junruren.com/posts/2025/10/MIT-CS-AI-Engagement</id><content type="html" xml:base="https://junruren.com/posts/2025/10/mit-cs-ai-engagement/"><![CDATA[<p>How do you build real connections <strong>outside</strong> your home academic program or department, especially in a place like MIT? This post shares a simple playbook for doing exactly that, using my own path into the CS/AI community as a case study. The short version: find the right announcement surfaces, <strong>show up</strong> (even when you won’t understand everything), and keep the curiosity dial set to “loud.”</p>

<p class="notice">This post also lives on <a href="https://junruren.substack.com/p/mit-cs-ai-engagement?utm_source=junruren.com&amp;utm_medium=referral&amp;utm_campaign=mit-cs-ai-engagement">my Substack</a> — comment there, or subscribe to get future posts by email.</p>

<h2 id="case-study-setup-an-sms-identity-crisis">Case Study Set‑Up: An SM’s “Identity Crisis”</h2>

<p>When I first arrived at MIT as an <strong>EECS Master of Science (SM)</strong> and <strong>MBA dual-degree</strong> candidate, I had a brief identity crisis about where I fit.</p>

<p>Conversations with other students or faculty often went like this:</p>

<blockquote>
  <p><strong>Who’s your PI (principal investigator)?</strong><br />
I don’t have one. I’m a master’s student.<br />
<strong>Oh, so you’re an MEng student then?</strong><br />
No, I’m SM 😅.</p>
</blockquote>

<p>This happens because <strong>the majority of EECS graduate students are PhD students</strong>, followed by MIT undergrads continuing for their MEng. That makes being an EECS SM both rare and meaningful. In my case, I cleared multiple admission committees (three, in fact: <strong>EECS, Sloan MBA, and LGO</strong>) 😎.</p>

<p>With no lab affiliation, no RA-ship, and no matched thesis advisor (well, now I do 😉), I still wanted to engage with the <strong>CS and AI research communities at MIT</strong>.</p>

<p><em>I want to be in the room where it happens.</em></p>

<p><img src="https://media2.giphy.com/media/v1.Y2lkPTc5MGI3NjExdm5jYjdxZm96OGd3ajFkZG03YmY1aXhnZGlkNzU2dm9mdTVjcnRoNiZlcD12MV9pbnRlcm5hbF9naWZfYnlfaWQmY3Q9Zw/3CYopbqdiUv4xJf7T6/giphy.gif" alt="" /></p>

<hr />

<h2 id="the-playbook-how-to-network-outside-your-home-program">The Playbook: How to Network Outside Your Home Program</h2>

<h3 id="1-set-one-clear-objective">1. Set one clear objective</h3>

<p>I decided one of my main grad school objectives would be to <strong>attend as many research seminars, talks, and guest lectures as possible</strong>. That single choice simplified a lot of decisions later (see: calendar collisions and FOMO).</p>

<h3 id="2-find-the-room-where-it-announces">2. Find the “Room Where It <em>Announces</em>”</h3>

<p>Every community has places where information lands first: department‑wide digests, lab/center mailing lists, seminar calendars, student collectives, Slack/Discord, and newsletters.</p>

<p>Your job is to <strong>discover and subscribe</strong>. (My MIT‑specific examples are in the appendix.)</p>

<h3 id="3-go-to-the-rooms-where-it-happens">3. Go to the Rooms Where It Happens</h3>

<p>Zoom is useful, but the hallway chat, the post‑talk question, and the walk‑and‑talk to the elevator are where many connections start.</p>

<p>In my first fall semester, I <em>sometimes</em> skipped a class (yes, one with attendance) to catch a talk.</p>

<p>Make your own call, but don’t sleep on the compounding value of being in the room.</p>

<h3 id="4-give-yourself-permission-not-to-understand-everything">4. Give yourself permission <strong>not</strong> to understand everything</h3>
<p>Do I always understand the talks? Absolutely not.</p>

<p>But every time, I leave with something new: a topic to Google later, a name to follow, a feel for where the field is heading. Even if I can’t follow the proofs, I can still catch the pulse, and that alone is worth being in the room.</p>

<p>For example, in one seminar, I heard <a href="https://arxiv.org/abs/2407.06460">“Machine <strong>Unlearning</strong>”</a> for the first time! That’s mind-opening when I was overwhelmingly exposed to “learning”.</p>

<hr />

<p>So if you also feel “between the labs,” don’t wait for an official invitation. Just show up. Sit in. Be curious.</p>

<p>If you’ve found other ways to engage (clubs, reading groups, random Slack channels, or secret pizza seminars) I’d love to hear them.<br />
<a href="mailto:junru@computer.org">Drop me a note</a>, and maybe I’ll see you in <strong>one of those rooms where it happens</strong>.</p>

<hr />

<h2 id="appendix-mit-csai-links-examples">Appendix: MIT CS/AI Links (Examples)</h2>

<p class="notice"><strong>Disclaimer.</strong> All links below point to <strong>public pages</strong> already available via CSAIL’s TIG website or public org pages. I’m <strong>re‑summarizing</strong> them here as examples, not advocating misuse. Please be respectful of list policies and norms. If a CSAIL administrator would prefer these subscription links not be referenced here, email me and I’ll remove them promptly.</p>

<ul>
  <li><strong>CSAIL (MIT Computer Science &amp; Artificial Intelligence Laboratory)</strong>
    <ul>
      <li>CSAIL homepage: <a href="https://www.csail.mit.edu/">https://www.csail.mit.edu/</a></li>
      <li>TIG: Popular Lab Mailing Lists: <a href="https://tig.csail.mit.edu/email-communicating/popular-lab-mailing-lists/">https://tig.csail.mit.edu/email-communicating/popular-lab-mailing-lists/</a></li>
      <li>Note: some lists (e.g., <code class="language-plaintext highlighter-rouge">csail-all</code>, <code class="language-plaintext highlighter-rouge">csail-internal</code>) are restricted.</li>
    </ul>
  </li>
  <li><strong>Student &amp; Research Collectives</strong>
    <ul>
      <li>Scale ML (cross‑lab MIT AI student collective): <a href="https://scale-ml.org/">https://scale-ml.org/</a>.
        <ul>
          <li>Just last month, they hosted <a href="https://x.com/chhillee"><strong>Horace He</strong></a>, who co-authored <a href="https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/"><em>“Defeating Nondeterminism in LLM Inference”</em></a>, the very first blog post from Thinking Machines.</li>
        </ul>
      </li>
      <li>ML Tea (informal machine learning talks with tea and usually snacks 😋): <a href="https://projects.csail.mit.edu/ml-tea/">https://projects.csail.mit.edu/ml-tea/</a></li>
      <li>MIT NLP Meetings Seminar Series: <a href="https://mitnlp.notion.site/">https://mitnlp.notion.site/</a></li>
    </ul>
  </li>
  <li><strong>Context</strong>
    <ul>
      <li>MIT EECS Graduate Programs - Admission Process (why SM/MEng confusion happens):<br />
<a href="https://www.eecs.mit.edu/academics/graduate-programs/admission-process/">https://www.eecs.mit.edu/academics/graduate-programs/admission-process/</a></li>
    </ul>
  </li>
</ul>]]></content><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><category term="mit" /><category term="networking" /><category term="grad-school" /><summary type="html"><![CDATA[How do you build real connections outside your home academic program or department, especially in a place like MIT? This post shares a simple playbook for doing exactly that, using my own path into the CS/AI community as a case study. The short version: find the right announcement surfaces, show up (even when you won’t understand everything), and keep the curiosity dial set to “loud.”]]></summary></entry><entry><title type="html">Cheatsheet for Your MIT Sloan Classes</title><link href="https://junruren.com/posts/2025/09/Cheatsheets/" rel="alternate" type="text/html" title="Cheatsheet for Your MIT Sloan Classes" /><published>2025-09-29T00:00:00+00:00</published><updated>2025-09-29T00:00:00+00:00</updated><id>https://junruren.com/posts/2025/09/MIT-Sloan-Cheatsheets</id><content type="html" xml:base="https://junruren.com/posts/2025/09/Cheatsheets/"><![CDATA[<p>Hi there. If you wind up on this page, I assume you’re busy studying for your upcoming MBA Core exams at MIT Sloan.</p>

<p>I was in your shoes in Fall 2024.</p>

<p>Originally, I just wanted to refamiliarize myself with LaTeX by typing up my study notes (instead of going to all the MBA parties, alas). Prior to MIT, I’d had success condensing everything into a single sheet of LaTeX, so I decided to do the same again; this time open-sourcing it for my fellow classmates.</p>

<p>So there you have it.</p>

<p class="notice"><strong>Caution</strong>: Your exam scope may differ from mine.</p>

<h2 id="mba-core">MBA Core</h2>

<ul>
  <li>15.010 Economic Analysis for Business Decisions (Fall 2024)
    <ul>
      <li><a href="https://github.com/junruren/MIT-Classes/blob/main/15.010/15_010_Midterm_Cheatsheet.pdf">Midterm</a></li>
      <li><a href="https://github.com/junruren/MIT-Classes/blob/main/15.010/15_010_Final_Cheatsheet.pdf">Final</a></li>
    </ul>
  </li>
  <li>15.515 Financial Accounting (Fall 2024)
    <ul>
      <li><a href="https://github.com/junruren/MIT-Classes/blob/main/15.515/15_515_H1_Cheatsheet.pdf">Midterm</a> — I didn’t make this one myself, so kudos to a legendary predecessor (please reach out if you made it!)</li>
      <li><a href="https://github.com/junruren/MIT-Classes/blob/main/15.515/15_515_Final_Cheatsheet.pdf">Final</a></li>
    </ul>
  </li>
</ul>

<h2 id="mba-core-elective">MBA Core Elective</h2>

<ul>
  <li>15.401 Managerial Finance (Spring 2025)
    <ul>
      <li><a href="https://github.com/junruren/MIT-Classes/blob/main/15.401/15_401_Midterm_Cheatsheet.pdf">Midterm</a></li>
      <li><a href="https://github.com/junruren/MIT-Classes/blob/main/15.401/15_401_Final_Cheatsheet.pdf">Final</a></li>
    </ul>
  </li>
</ul>

<h2 id="lgo-core">LGO Core</h2>

<ul>
  <li>15.086 Engineering Probability (Summer 2024)
    <ul>
      <li><a href="https://github.com/junruren/MIT-Classes/blob/main/15.086/15_086_Cheatsheet.pdf">Final</a></li>
    </ul>
  </li>
  <li>15.087 Engineering Statistics and Data Science (Summer 2024)
    <ul>
      <li><a href="https://github.com/junruren/MIT-Classes/blob/main/15.087/15_087_Cheatsheet.pdf">Final</a></li>
    </ul>
  </li>
</ul>

<hr />

<h2 id="tips-on-printing">Tips on Printing</h2>

<ol>
  <li>Ideally, find a color printer.</li>
  <li>Select the “Shrink to Fit” option in your printer settings.</li>
  <li>Select the highest possible print quality.</li>
</ol>

<p><em>Good luck!</em></p>

<hr />

<p>The best way to send me a “thank you note” is to <strong>🌟 Star</strong> my <a href="https://github.com/junruren/MIT-Classes/">GitHub repository</a> that hosts all these cheatsheets 😉 so… my thanks to you!</p>]]></content><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><category term="mitsloan" /><category term="mit" /><category term="class" /><summary type="html"><![CDATA[Hi there. If you wind up on this page, I assume you’re busy studying for your upcoming MBA Core exams at MIT Sloan.]]></summary></entry><entry><title type="html">Tutorial: MIT Thesis LaTeX Template with VS Code</title><link href="https://junruren.com/posts/2025/08/MIT-Thesis-LaTeX/" rel="alternate" type="text/html" title="Tutorial: MIT Thesis LaTeX Template with VS Code" /><published>2025-08-04T00:00:00+00:00</published><updated>2025-08-04T00:00:00+00:00</updated><id>https://junruren.com/posts/2025/08/Tutorial-MIT-Thesis-LaTeX</id><content type="html" xml:base="https://junruren.com/posts/2025/08/MIT-Thesis-LaTeX/"><![CDATA[<p>Let’s set up the MIT thesis LaTeX template locally!</p>

<p><img src="/images/2025-08-04-Tutorial-MIT-Thesis-LaTeX/Cover_by_ChatGPT.png" alt="Cover Photo by ChatGPT" /></p>
<blockquote>
  <p>Generated by GPT-4o. Prompt: <em>A playful illustration showing a happy MIT student proudly holding a printed MIT thesis in one hand and a laptop showing VS Code with LaTeX code in the other. In the background, icons representing LaTeX, PDF documents, and MIT’s iconic Great Dome subtly float around. The student stands triumphantly on top of a stack of neatly bound theses. The style is colorful, lighthearted, and cartoonish.</em></p>
</blockquote>

<p>MIT’s libraries require theses to be deposited electronically using a strict format. To simplify formatting, the <a href="https://web.mit.edu/thesis/tex/"><strong>MIT thesis LaTeX template</strong></a> provides a class (<code class="language-plaintext highlighter-rouge">mitthesis.cls</code>) and a set of sample files that implement these requirements.</p>

<p>Many thanks to Prof. John H. Lienhard for maintaining <code class="language-plaintext highlighter-rouge">mitthesis</code> and for generously answering questions from students like me.</p>

<p class="notice">Update September 2026: this post now targets <a href="https://ctan.org/pkg/mitthesis"><code class="language-plaintext highlighter-rouge">mitthesis</code> v1.24</a>, dated August 21, 2026. If you read an earlier version, the one thing worth knowing is that <strong><code class="language-plaintext highlighter-rouge">\LGO</code> is now an official, documented command</strong>, so the temporary workarounds this post used to describe are no longer needed. Any other screenshots or comments you remember came from the v1.20-v1.23 period, when class-file location, committee-page behavior, and LGO cover-page workarounds were all still moving.</p>

<p>This tutorial blog builds on top of my previous <a href="/posts/2025/06/LaTeX-VSCode/">“Tutorial: Use LaTeX Locally with VS Code”</a>. By following that tutorial, you should already have:</p>

<ul>
  <li><strong>Visual Studio Code with LaTeX Workshop installed.</strong> The extension provides core LaTeX features such as auto‑building to PDF, an integrated PDF viewer, SyncTeX navigation, IntelliSense, and log parsing. It automatically runs the sequence of tools needed to build your document and highlights errors in the editor.</li>
  <li><strong>TeX Live 2022 or newer.</strong> The MIT thesis class requires a recent LaTeX distribution; <a href="https://ctan.org/pkg/mitthesis">CTAN lists v1.24</a> as requiring TeX Live 2022 or later. A full TeX Live installation includes <code class="language-plaintext highlighter-rouge">lualatex</code>, <code class="language-plaintext highlighter-rouge">pdflatex</code>, biber and other programs needed by the template. I recommend LuaLaTeX for the current template, especially if you care about modern PDF metadata or accessibility workflows.</li>
  <li><strong>Biber for bibliography management.</strong> The template defaults to using biblatex with the biber backend. Biber is part of TeX Live and will run automatically if configured in LaTeX Workshop.</li>
</ul>

<h2 id="1-acquire-the-template">1. Acquire the template</h2>

<p>Download it here: <a href="https://mirrors.ctan.org/macros/latex/contrib/mitthesis.zip">https://mirrors.ctan.org/macros/latex/contrib/mitthesis.zip</a>.</p>

<p class="notice">Comprehensive TeX Archive Network (CTAN) is a network of websites that serve as the central repository for TeX-related materials and software; it’s legit and referenced by the official  <a href="https://web.mit.edu/thesis/tex/">MIT thesis LaTeX template</a> website.</p>

<p>Extract (i.e., “unzip”) the archive to a convenient location—ideally the folder where you plan to write your thesis.</p>

<p><img src="/images/2025-08-04-Tutorial-MIT-Thesis-LaTeX/mitthesis-unzipped.png" alt="Example view of the unzipped `mitthesis`" /></p>

<h3 id="understand-the-contents">Understand the contents</h3>

<p>It’s always recommended to read the included <code class="language-plaintext highlighter-rouge">README.md</code>. This file lists the archive contents, notably a class file and a MIT-thesis-template folder containing everything you need to start writing:</p>

<table>
  <thead>
    <tr>
      <th>File/Folder</th>
      <th>Purpose</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">mitthesis.cls</code></td>
      <td>Core LaTeX class implementing MIT formatting</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">MIT-thesis-template/MIT-Thesis.tex</code></td>
      <td>Main LaTeX file for your thesis</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">abstract.tex</code>, <code class="language-plaintext highlighter-rouge">acknowledgments.tex</code>, <code class="language-plaintext highlighter-rouge">biosketch.tex</code></td>
      <td>Files where you insert your abstract, acknowledgments and optional biographical sketch</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">chapter1.tex</code>, <code class="language-plaintext highlighter-rouge">chapter...</code></td>
      <td>Sample chapters to use as templates</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">\Reader{...}</code> commands / <code class="language-plaintext highlighter-rouge">committee_members.tex</code></td>
      <td>In v1.21 and newer, <code class="language-plaintext highlighter-rouge">\Reader{...}</code> commands automatically generate the thesis committee page; if you omit all readers, you can still insert your own optional <code class="language-plaintext highlighter-rouge">committee_members.tex</code> page before the abstract</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">appendixa.tex</code>, <code class="language-plaintext highlighter-rouge">appendixb.tex</code></td>
      <td>Sample appendices showing code listing and long tables</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">mitthesis-sample.bib</code></td>
      <td>Sample bibliography file with many entries</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">mitthesis-style.css</code></td>
      <td>Optional CSS embedded when tagged PDF is in use</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">mydesign.tex</code></td>
      <td>Optional file where you can load packages to customise colours, margins or caption styles</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">fontsets/</code></td>
      <td>Subdirectory containing optional font definitions</td>
    </tr>
  </tbody>
</table>

<p>Additionally, the <code class="language-plaintext highlighter-rouge">mitthesis-doc</code> directory contains detailed PDF documentation, and the examples directory provides sample theses showcasing different font options.</p>

<p>After extraction, keep the directory structure intact; LaTeX will look for chapter files relative to the main file. You can rename the outer folder to reflect your project’s name.</p>

<h3 id="class-file-location-update">Class file location update</h3>

<p>Update September 2026: the current <a href="https://ctan.org/pkg/mitthesis"><code class="language-plaintext highlighter-rouge">mitthesis</code> package listing</a> shows <code class="language-plaintext highlighter-rouge">mitthesis</code> v1.24 dated August 21, 2026. The archive still puts <code class="language-plaintext highlighter-rouge">mitthesis.cls</code> in the outer <code class="language-plaintext highlighter-rouge">mitthesis</code> folder, while <code class="language-plaintext highlighter-rouge">MIT-thesis-template</code> is the folder with the files you edit. The <a href="https://mirrors.ctan.org/macros/latex/contrib/mitthesis/mitthesis-doc/mitthesis-doc.pdf">official documentation</a> says to copy <code class="language-plaintext highlighter-rouge">MIT-thesis-template</code> onto your system; if the current <code class="language-plaintext highlighter-rouge">mitthesis.cls</code> is already installed in TeX Live, you are all set, and if not, copy <code class="language-plaintext highlighter-rouge">mitthesis.cls</code> into your working directory.</p>

<p>In practice, with VS Code + LaTeX Workshop, the least surprising setup is:</p>

<ol>
  <li>Open <code class="language-plaintext highlighter-rouge">MIT-thesis-template</code> as the VS Code workspace.</li>
  <li>Copy <code class="language-plaintext highlighter-rouge">../mitthesis.cls</code> into <code class="language-plaintext highlighter-rouge">MIT-thesis-template</code> if your local TeX Live has an older installed class.</li>
  <li>Rebuild.</li>
</ol>

<p class="notice">This matters because TeX Live may already resolve <code class="language-plaintext highlighter-rouge">\documentclass{mitthesis}</code> to an older installed <code class="language-plaintext highlighter-rouge">mitthesis.cls</code>. For example, a TeX Live 2025 installation can still have mitthesis v1.20 while the current CTAN zip is v1.24. Mixing template files and class files from different versions can trigger confusing errors; <code class="language-plaintext highlighter-rouge">Undefined control sequence \CiteNolink</code> is one example. Copying the outer class file into <code class="language-plaintext highlighter-rouge">MIT-thesis-template</code> makes the project use the class version that came with the files you just downloaded.</p>

<h2 id="2-opening-the-project-in-vs-code">2. Opening the project in VS Code</h2>

<p>Launch <strong>VS Code</strong>, then choose <strong>File → Open Folder…</strong> and select the <code class="language-plaintext highlighter-rouge">MIT-thesis-template</code> folder (not the outer <code class="language-plaintext highlighter-rouge">mitthesis</code> folder!).</p>

<p>VS Code will treat this folder as the workspace.</p>

<p>Open the <code class="language-plaintext highlighter-rouge">MIT-Thesis.tex</code> file as it is the root document.</p>

<p><img src="/images/2025-08-04-Tutorial-MIT-Thesis-LaTeX/MIT-thesis-template-VSCode.png" alt="MIT-thesis-template folder opened in VS Code" /></p>

<h2 id="3-trigger-the-build">3. Trigger the build</h2>

<p>You should see a “TEX” badge from the LaTeX Workshop extension appearing in the leftmost panel of your window.</p>

<p>Trigger a file save using standard shortcuts:</p>
<ul>
  <li>macOS: <code class="language-plaintext highlighter-rouge">command</code> + <code class="language-plaintext highlighter-rouge">S</code></li>
  <li>Windows: <code class="language-plaintext highlighter-rouge">Ctrl</code> + <code class="language-plaintext highlighter-rouge">S</code></li>
</ul>

<p>As explained in the <a href="/posts/2025/06/LaTeX-VSCode/">previous tutorial</a>, saving the file automatically triggers the LaTeX Workshop build process. You’ll notice a spinning “🔄 Build” icon in the bottom-left corner, indicating compilation in progress. This may take a minute.</p>

<p class="notice">During compilation, local setup variations can cause errors. If you encounter an issue, please let me know so I can include quick fixes here. I also recommend troubleshooting with your preferred GenAI tool—my go-to choice is GitHub Copilot.</p>

<p><img src="/images/2025-08-04-Tutorial-MIT-Thesis-LaTeX/MIT-thesis-first-compiled-VSCode.png" alt="Default MIT-Thesis compiled with resulting PDF displayed" /></p>

<p>Scroll through the generated PDF file, <code class="language-plaintext highlighter-rouge">MIT-Thesis.pdf</code>; it should closely match the <a href="http://mirrors.ctan.org/macros/latex/contrib/mitthesis/examples/font_samples/Lmodern_sample.pdf">official example PDF</a>.</p>

<p>Congratulations, your MIT Thesis LaTeX template is now ready to use.</p>

<h2 id="4-familiarize-yourself-with-the-template">4. Familiarize yourself with the template</h2>

<p>The best way to get comfortable is to explore its structure – no shortcuts here.</p>

<p><strong>Please read the <a href="http://mirrors.ctan.org/macros/latex/contrib/mitthesis/mitthesis-doc/mitthesis-doc.pdf">official documentation</a> – again, no shortcut</strong>.</p>

<p>One quick tip: skim through <code class="language-plaintext highlighter-rouge">MIT-Thesis.tex</code>. Notice the multiple <code class="language-plaintext highlighter-rouge">\input{}</code> statements pulling in separate files for different thesis sections, such as <code class="language-plaintext highlighter-rouge">abstract.tex</code> and <code class="language-plaintext highlighter-rouge">chapter1.tex</code>. Edit these files slightly, rebuild the PDF, and observe the results – action learning at its finest!</p>

<p>If you are a student in MIT’s <a href="https://lgo.mit.edu/">Leaders for Global Operations (LGO)</a> dual-degree program, continue to the next section. Otherwise, you’re all set.</p>

<h2 id="5-lgo-thesis-tweaks">5. LGO Thesis tweaks</h2>

<p>The official package includes a useful dual-degree example: <a href="https://mirrors.ctan.org/macros/latex/contrib/mitthesis/examples/cover_page_samples/latex_sources/One_author_two_degrees.tex"><code class="language-plaintext highlighter-rouge">One_author_two_degrees.tex</code></a>. The LGO-specific version below follows the current v1.24 interface, where the LGO cover-page line is built in.</p>

<p class="notice">Before the mechanics: <strong>thank you to <a href="https://lienhard.mit.edu/people/#lienhard">Prof. John H. Lienhard</a></strong> for maintaining <code class="language-plaintext highlighter-rouge">mitthesis</code> and for the generous email exchange behind this section. It started in May 2026 with a small bug report from me (the sample dual-degree file omitted “May” from its list of valid degree months), which he fixed the same morning, along with pointing me to the cleaner <code class="language-plaintext highlighter-rouge">\\ &amp;</code> pattern for multi-department lines. When I asked whether the package might natively support the LGO cover-page phrase, he offered to write a <code class="language-plaintext highlighter-rouge">\LGO</code> macro, checked it with MIT Libraries, and kept me posted through every step. He also wrote up the whole modernization effort in <em>TUGboat</em>: <a href="https://dspace.mit.edu/handle/1721.1/173917">“Modernizing MIT’s thesis template: <code class="language-plaintext highlighter-rouge">mitthesis.cls</code>”</a>, TUGboat 47(2), 2026, pp. 212-221 (<a href="https://doi.org/10.47397/tb/47-2/tb146lienhard-mitthesis">doi:10.47397/tb/47-2/tb146lienhard-mitthesis</a>). It is a good read even if you never touch a class file.</p>

<ol>
  <li>In your <code class="language-plaintext highlighter-rouge">MIT-Thesis.tex</code>, locate:
    <div class="language-tex highlighter-rouge"><div class="highlight"><pre class="highlight"><code> <span class="nt">\begin{document}</span>
 <span class="c">%%% edit the following commands to match your thesis %%%%%%%%%%</span>
</code></pre></div>    </div>
  </li>
  <li>Replace the content from <code class="language-plaintext highlighter-rouge">\title{...}</code> to <code class="language-plaintext highlighter-rouge">\ThesisDate{...}</code> (inclusive) with this:
    <div class="language-tex highlighter-rouge"><div class="highlight"><pre class="highlight"><code> <span class="k">\title</span><span class="p">{</span>Simplify and Accelerate: An Awesome Dual-Degree Thesis Title<span class="p">}</span>

 <span class="c">% \Author{Author full name}{Author department}[Author's first PREVIOUS degree][Author's second PREVIOUS degree][...</span>
 <span class="c">% Note that third, fourth, fifth, and sixth arguments are optional [] and may be omitted</span>

 <span class="k">\Author</span><span class="p">{</span>LGO Student Name<span class="p">}{</span>MIT Sloan School of Management and <span class="k">\\</span> <span class="p">&amp;</span> Department of Electrical Engineering and Computer Science<span class="p">}</span>[B.S. Previous Degree, Previous College, 2018]

 <span class="c">% Use once for each degree fulfilled by thesis</span>
 <span class="c">% For two degrees from one department, leave the department argument blank for the second degree {}.</span>
 <span class="k">\Degree</span><span class="p">{</span>Master of Business Administration<span class="p">}{</span>MIT Sloan School of Management<span class="p">}</span>
 <span class="k">\Degree</span><span class="p">{</span>Master of Science in Electrical Engineering and Computer Science<span class="p">}{</span>Department of Electrical Engineering and Computer Science<span class="p">}</span>

 <span class="c">% If there is more than one supervisor, use the \Supervisor command for each.</span>
 <span class="c">% The optional [department] field is used on the automatically generated thesis committee page.</span>
 <span class="k">\Supervisor</span><span class="p">{</span>Sloan Advisor Name<span class="p">}{</span>Professor of Operations Management<span class="p">}</span>[MIT Sloan School of Management]
 <span class="k">\Supervisor</span><span class="p">{</span>Engineering Advisor Name<span class="p">}{</span>Professor of Electrical Engineering and Computer Science<span class="p">}</span>[Department of Electrical Engineering and Computer Science]

 <span class="c">% Professor who formally accepts theses for your department (e.g., the Graduate Officer, Professor Sméagol,...)</span>
 <span class="c">% If you need to reduce vertical space, put the acceptor title in the second argument and leave the third blank {}.</span>
 <span class="k">\Acceptor</span><span class="p">{</span>Engineering Acceptor Name<span class="p">}{</span>Professor of Electrical Engineering and Computer Science<span class="p">}{</span>Graduate Officer, Department of Electrical Engineering and Computer Science<span class="p">}</span>
 <span class="k">\Acceptor</span><span class="p">{</span>Sloan Acceptor Name<span class="p">}{</span>Assistant Dean<span class="p">}{</span>MBA Program, Sloan School of Management<span class="p">}</span>

 <span class="c">% Keep the signature block at normal size unless the title page itself needs more vertical space.</span>
 <span class="k">\SignatureBlockSize</span><span class="p">{</span><span class="k">\normalsize</span><span class="p">}</span>
 <span class="k">\AuthorNameSize</span><span class="p">{</span><span class="k">\normalsize</span><span class="p">}</span>
 <span class="k">\Tighten</span>

 <span class="c">% Usage: \DegreeDate{Month}{year}</span>
 <span class="k">\DegreeDate</span><span class="p">{</span>May<span class="p">}{</span>2026<span class="p">}</span>

 <span class="c">% Date that final thesis is submitted to department</span>
 <span class="k">\ThesisDate</span><span class="p">{</span>May 8, 2026<span class="p">}</span>
</code></pre></div>    </div>
  </li>
</ol>

<p>The important part for long department names is the <code class="language-plaintext highlighter-rouge">\\ &amp;</code> line break inside the second argument to <code class="language-plaintext highlighter-rouge">\Author</code>. The current documentation uses this pattern for multiple departments and titles. An earlier version of this blog (probably referred to by my Class of 2026 classmates) recommended <code class="language-plaintext highlighter-rouge">\shortstack[l]{...\\...}</code>; that worked visually, but the <code class="language-plaintext highlighter-rouge">\\ &amp;</code> style now matches the template’s own examples better.</p>

<p>Per “Thesis Review and Submission Process”, <em>LGO Handbook</em> (accessed on August 3, 2025), we also need to include:</p>

<blockquote>
  <p>“IN CONJUNCTION WITH THE LEADERS FOR GLOBAL OPERATIONS PROGRAM AT THE MASSACHUSETTS INSTITUTE OF TECHNOLOGY”.</p>
</blockquote>

<p>Good news: the template already includes the command you need. In <code class="language-plaintext highlighter-rouge">MIT-Thesis.tex</code>, scroll past the <code class="language-plaintext highlighter-rouge">\Reader{...}</code> commands and the Creative Commons block to this banner:</p>

<div class="language-tex highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c">%%%%%%  Special additions to title page  %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%</span>

<span class="c">% A few departments or programs ask students to add special text to the title page:</span>

<span class="c">% \LGO % uncomment this command if you are in the Leaders for Global Operations program</span>
<span class="c">% \TitlePageProgramNote{...program's required text...} % use for other programs that require added title page text</span>

<span class="c">% &gt;&gt;&gt; DO NOT use these commands unless your program requires it &lt;&lt;&lt;</span>
</code></pre></div></div>

<p><strong>Uncomment one line.</strong> Delete the <code class="language-plaintext highlighter-rouge">%</code> in front of <code class="language-plaintext highlighter-rouge">\LGO</code>:</p>

<div class="language-tex highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">\LGO</span> <span class="c">% uncomment this command if you are in the Leaders for Global Operations program</span>
</code></pre></div></div>

<p>That is the entire LGO tweak. (The official documentation says to put <code class="language-plaintext highlighter-rouge">\LGO</code> “in the preamble”; in the template the line sits after <code class="language-plaintext highlighter-rouge">\begin{document}</code>, alongside <code class="language-plaintext highlighter-rouge">\title</code> and <code class="language-plaintext highlighter-rouge">\ThesisDate</code>. Either works, as long as it comes before <code class="language-plaintext highlighter-rouge">\maketitle</code>.)</p>

<p>The <a href="http://mirrors.ctan.org/macros/latex/contrib/mitthesis/mitthesis-doc/mitthesis-doc.pdf">official documentation</a> covers this in §7.3, “Adding text to title page when required by your program,” which names LGO as its worked example. If you are in a different program that requires its own line, <code class="language-plaintext highlighter-rouge">\TitlePageProgramNote{your text}</code> is the general-purpose version. The documentation attaches a caution worth repeating verbatim:</p>

<blockquote>
  <p>Do not add text unless your program has approval for the addition from the MIT Libraries; otherwise, your thesis may not be accepted.</p>
</blockquote>

<p>LGO has that approval, which is why the command exists.</p>

<p class="notice">Using Overleaf instead of a local install? As of September 2026, Prof. Lienhard notes that Overleaf’s template gallery is still on <strong>v1.21</strong> (dated November 2, 2025), which does not have <code class="language-plaintext highlighter-rouge">\LGO</code> at all, so uncommenting it there fails with <code class="language-plaintext highlighter-rouge">Undefined control sequence</code>. Until the gallery catches up, upload the current CTAN <code class="language-plaintext highlighter-rouge">mitthesis.cls</code> into your Overleaf project. Overleaf also carries a third-party <a href="https://www.overleaf.com/latex/templates/lgo-thesis-template/txmvvktbdxst">LGO Thesis Template</a>, described as a slight modification of the MIT one for LGO fellows; it is not Prof. Lienhard’s official package and lags the CTAN version, so check which class file it actually contains before trusting it.</p>

<p class="notice">Breadcrumb for my Class of 2026 classmates: earlier versions of this post walked through hand-rolled ways to get the LGO line onto the cover page, back when <code class="language-plaintext highlighter-rouge">\LGO</code> was not yet active. If you already submitted using one of them, there is nothing to revisit; your cover page was correct. If you are rebuilding on v1.24, delete any leftover <code class="language-plaintext highlighter-rouge">\ExplSyntaxOn</code> block from your preamble, or <code class="language-plaintext highlighter-rouge">\LGO</code> gets defined twice and LaTeX will error.</p>

<p>Rebuild your LaTeX project and you should see a cover page like this:</p>

<p><img src="/images/2025-08-04-Tutorial-MIT-Thesis-LaTeX/LGO-Thesis-Cover-Example.jpg" alt="An example of LGO thesis cover page rendered" /></p>

<p class="notice"><strong>Note</strong> that this tutorial uses <strong>May 2026</strong> for the LGO Class of 2026 cover page. MIT supports February, May, June, and September as degree months in the template. Requirements do get revised, so confirm your degree date and the rest of the title page with your department and the LGO program office before submitting.</p>

<hr />

<p>Good luck with your thesis! Good luck to me too…</p>]]></content><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><category term="tutorial" /><category term="latex" /><category term="vscode" /><category term="mit" /><summary type="html"><![CDATA[Let’s set up the MIT thesis LaTeX template locally!]]></summary></entry><entry><title type="html">Tutorial: Use LaTeX Locally with VS Code</title><link href="https://junruren.com/posts/2025/06/LaTeX-VSCode/" rel="alternate" type="text/html" title="Tutorial: Use LaTeX Locally with VS Code" /><published>2025-06-29T00:00:00+00:00</published><updated>2025-06-29T00:00:00+00:00</updated><id>https://junruren.com/posts/2025/06/Tutorial-LaTeX-VSCode</id><content type="html" xml:base="https://junruren.com/posts/2025/06/LaTeX-VSCode/"><![CDATA[<p>This post is a quick guide to setting up LaTeX on your own computer, so you can write papers, resumes, or <a href="https://github.com/junruren/MIT-Classes">dazzling “class notes”</a> using LaTeX <strong>locally</strong>—without relying on internet connectivity. In other words, you won’t be tied to <a href="https://www.overleaf.com/">Overleaf</a>: you can bring your laptop anywhere and keep working on your next big ideas.</p>

<h2 id="why-not-overleaf">Why not Overleaf?</h2>

<p>First off, let me say that <a href="https://www.overleaf.com/">Overleaf</a> is awesome! It’s a rather complete product, crafted for all the heavy LaTeX writers, with many desirable features, a huge library of <a href="https://www.overleaf.com/latex/templates">templates</a>, and a great collection of <a href="https://www.overleaf.com/learn">LaTeX guides</a>.</p>

<p><img src="/images/2025-06-29-Tutorial-LaTeX-VSCode/Cover_by_ChatGPT.png" alt="Cover Photo by ChatGPT" /></p>
<blockquote>
  <p>Generated by GPT-4o. Prompt: <em>A playful illustration showing a person sitting on an airplane, trying to work on a laptop with the Overleaf logo on the screen, but looking frustrated because there’s no Wi-Fi signal. Outside the window, fluffy clouds and a blue sky. In the background, a thought bubble shows the same person happily working on their laptop (with VS Code logo) in a cozy home office, surrounded by books and coffee, with a checkmark and a glowing PDF icon. The style is colorful, lighthearted, and cartoonish.</em></p>
</blockquote>

<p>However, relying solely on a cloud-based service like Overleaf comes with a few drawbacks:</p>

<h3 id="1-server-downtime">1. Server downtime</h3>

<p>Just like when your favorite social media or streaming site occasionally refuses to load, Overleaf is subject to outages. See <a href="https://status.overleaf.com/history">Overleaf’s status page</a> for yourself.</p>

<p>One of my MIT LGO alumni friends mentioned that a few years ago, while writing their MIT thesis, an Overleaf outage definitely made for some uneasy times.</p>

<p>At the time of writing this blog, the most recent “blockbuster” incident was Overleaf being down on May 14, just before the <a href="https://x.com/DianboLiu/status/1922544766849257851">NeurIPS manuscript deadline</a>. You can see how people reacted under <a href="https://x.com/overleaf/status/1922576431130759359">Overleaf’s X post about this outage</a>.</p>

<h3 id="2-difficult-to-write-without-internet-eg-onboard-a-flight">2. Difficult to write without internet (e.g., onboard a flight)</h3>

<p>Ideas come and go. Who doesn’t like the idea of being able to jot them down and immediately see the LaTeX render? Many people seem to be more productive in the air, but are we going to pay for that in-flight Wi-Fi (assuming Wi-Fi is even available onboard)?</p>

<p>This is where a local setup shines. While Overleaf is truly a crucial product in academia and beyond (and I still like it!), I also want to be immune to these risks. Here’s how you can de-risk by having a local setup.</p>

<hr />

<p class="notice"><strong>Disclaimer:</strong> I am a heavy macOS user, so for the rest of this guide, I’ll do my best to mention what might be different on Windows. (If you use Unix, this guide <em>may</em> already not be needed for you!)</p>

<h2 id="tldr">TL;DR</h2>

<p>My local setup is: VS Code + the “LaTeX Workshop” extension for VS Code + TeX Live.</p>

<p>This blog is aggressively simplified from the much more thorough <a href="https://github.com/James-Yu/LaTeX-Workshop/wiki/Install">installation guide of the LaTeX Workshop extension</a>. I envisioned this blog to be very lightweight and introductory (i.e., not loaded with info that would please “power users”), but I’ll iterate and try to find the right balance between too basic and too hardcore. Your feedback is greatly appreciated!</p>

<h2 id="1-install-tex-live">1. Install TeX Live</h2>

<p><a href="https://www.tug.org/texlive/">TeX Live</a> is a LaTeX distribution compatible with and recommended by the LaTeX Workshop extension.</p>

<ul>
  <li>macOS: <a href="https://www.tug.org/mactex/mactex-download.html">https://www.tug.org/mactex/mactex-download.html</a></li>
  <li>Windows: <a href="https://www.tug.org/texlive/windows.html#install">https://www.tug.org/texlive/windows.html#install</a></li>
</ul>

<h2 id="2-install-vs-code">2. Install VS Code</h2>

<p>Download it here: <a href="https://code.visualstudio.com/download">https://code.visualstudio.com/download</a></p>

<h2 id="3-install-the-latex-workshop-extension-for-vs-code">3. Install the LaTeX Workshop Extension for VS Code</h2>

<p>Open VS Code and look up the “LaTeX Workshop” extension in the Extension Marketplace:</p>

<ol>
  <li>Bring up the Extensions view by clicking on the Extensions icon in the Activity Bar on the side of VS Code.</li>
  <li>In the search bar, type “LaTeX Workshop” and look for the one authored by James Yu.</li>
  <li>Click “Install”.</li>
</ol>

<p><img src="/images/2025-06-29-Tutorial-LaTeX-VSCode/LaTeXWorkshop-Extension-VSCode-Screenshot.png" alt="Screenshot of LaTeX Workshop shown in VS Code's Extension Marketplace" /></p>

<h2 id="4-test-your-first-locally-built-pdf">4. Test: Your first locally built PDF</h2>

<p>Let’s try it out! In VS Code, create a new file. I named mine <code class="language-plaintext highlighter-rouge">hello.tex</code> and typed in the following content (<a href="https://www.overleaf.com/learn/latex/Learn_LaTeX_in_30_minutes#Writing_your_first_piece_of_LaTeX">source: Overleaf guide</a>—see, I told you Overleaf is awesome):</p>

<div class="language-tex highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">\documentclass</span><span class="p">{</span>article<span class="p">}</span>
<span class="nt">\begin{document}</span>
First document. This is a simple example, with no 
extra parameters or packages included.
<span class="nt">\end{document}</span>
</code></pre></div></div>

<p><img src="/images/2025-06-29-Tutorial-LaTeX-VSCode/VSCode-First-LaTeX-Before-Build.png" alt="Screenshot of a simple tex file created on VS Code before saving" /></p>

<p>Now, save your file in VS Code:</p>
<ul>
  <li>macOS: <code class="language-plaintext highlighter-rouge">command</code> + <code class="language-plaintext highlighter-rouge">S</code></li>
  <li>Windows: <code class="language-plaintext highlighter-rouge">Ctrl</code> + <code class="language-plaintext highlighter-rouge">S</code></li>
</ul>

<p>By default, the LaTeX Workshop extension has a convenient feature turned on: whenever you save or even change the <code class="language-plaintext highlighter-rouge">.tex</code> file, compilation of your LaTeX project is triggered. If nothing seems to happen, try clicking the green triangle icon at the top right corner.</p>

<p class="notice">Like anything in the developer world, the LaTeX Workshop extension comes with many customization options. For example, the auto build trigger is available in the extension’s settings. <img src="/images/2025-06-29-Tutorial-LaTeX-VSCode/LaTeXWorkshop-AutoBuild.png" alt="Screenshot of LaTeX Workshop Setting on Auto Build" /> For simplicity, I won’t go into detail here.</p>

<p>You should now see something similar to the following screenshot:</p>

<p><img src="/images/2025-06-29-Tutorial-LaTeX-VSCode/VSCode-First-LaTeX-After-Build.png" alt="Screenshot of a simple tex file created on VS Code after saving" /></p>

<p>In the same folder as your <code class="language-plaintext highlighter-rouge">hello.tex</code>, you’ll see:</p>

<ul>
  <li>a <code class="language-plaintext highlighter-rouge">hello.pdf</code> (what we really care about), and</li>
  <li>a bunch of auxiliary files; ignore them for now (again, I won’t go into detail here).</li>
</ul>

<p>You can click on <code class="language-plaintext highlighter-rouge">hello.pdf</code> and expect to see exactly what our simple <code class="language-plaintext highlighter-rouge">.tex</code> file should render into.</p>

<p>A very convenient feature is to show the PDF file side-by-side with your <code class="language-plaintext highlighter-rouge">.tex</code> file. Find the green triangle Build button at the top right corner; right next to it is another icon with a little magnifier over two rectangles—click it.</p>

<p><img src="/images/2025-06-29-Tutorial-LaTeX-VSCode/LaTeX-Compiled-PDF-Example.png" alt="Screenshot of the simple tex file with its compiled PDF opened side-by-side" /></p>

<p>Congratulations! You now have a “barebones” setup for working on LaTeX documents locally. Convince yourself by disconnecting your computer from the internet, making some random edits to your <code class="language-plaintext highlighter-rouge">.tex</code> file, and saving your changes again.</p>

<h3 id="troubleshooting">Troubleshooting</h3>

<p>I initially came across <code class="language-plaintext highlighter-rouge">LaTeX fatal error on PID undefined. Error: spawn latexmk ENOENT</code>. My solution was simply to restart VS Code.</p>

<p>Please let me know if you encounter any other issues!</p>

<hr />

<h2 id="coming-soon">Coming soon…</h2>

<p>I plan to write a mini-series of blogs on LaTeX, specifically in the context of writing a thesis at MIT. In the future, I plan to cover:</p>

<ul>
  <li>How to use <code class="language-plaintext highlighter-rouge">git</code> to track your progress and back up with GitHub</li>
  <li><a href="/posts/2025/08/MIT-Thesis-LaTeX/">How to use the MIT thesis template with this local setup</a></li>
</ul>

<p>Meanwhile, please let me know if you have any suggestions for this tutorial or future topics.</p>]]></content><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><category term="tutorial" /><category term="latex" /><category term="vscode" /><summary type="html"><![CDATA[This post is a quick guide to setting up LaTeX on your own computer, so you can write papers, resumes, or dazzling “class notes” using LaTeX locally—without relying on internet connectivity. In other words, you won’t be tied to Overleaf: you can bring your laptop anywhere and keep working on your next big ideas.]]></summary></entry><entry><title type="html">6.8300 Final Project: Taming CLIP’s Captioning Bias: A COCO-Driven Analysis and Permutation Ensemble</title><link href="https://junruren.com/posts/2025/05/6.8300-final/" rel="alternate" type="text/html" title="6.8300 Final Project: Taming CLIP’s Captioning Bias: A COCO-Driven Analysis and Permutation Ensemble" /><published>2025-05-13T00:00:00+00:00</published><updated>2025-05-13T00:00:00+00:00</updated><id>https://junruren.com/posts/2025/05/6.8300-Final-Project-Blog</id><content type="html" xml:base="https://junruren.com/posts/2025/05/6.8300-final/"><![CDATA[<p><strong>Abstract</strong>: <em>Vision-language models like CLIP struggle with multi-object scenes, often favoring prominent objects or those mentioned first in captions. Using real-world COCO images, we show that CLIP’s caption-matching accuracy drops from 91.23% to 87.45% when object order is reversed. To address this, we explore a post-hoc mitigation: a permutation ensemble that averages scores across all object orders, boosting robustness and recovering accuracy to 90.04%. Our findings reveal persistent order biases and offer a simple, effective strategy to improve CLIP’s reliability in complex scenes.</em></p>

<p><img src="/images/2025-05-13-6.8300-Final-Project-Blog/Cover_by_ChatGPT.png" alt="Cover Photo by ChatGPT" /></p>

<blockquote>
  <p>Generated by GPT-4o. Prompt: <em>Please create a widescreen cover image illustration for my blog post. Please illustrate a “humanized” CLIP model being unsure about which caption best describes the attached image. The captions are: “a pizza and a dog and a dining table”, “a dining table and a dog and a pizza”，”a pizza and a dog and a cell phone”. Please include the provided image along with three captions in the generated image!</em></p>
</blockquote>

<hr />

<h2 id="introduction">Introduction</h2>

<p>Accurate image captioning, a key challenge at the intersection of computer vision and natural language processing, is fundamental for enabling machines to “see” and describe visual content. This capability underpins diverse applications, from accessibility tools that interpret images for visually impaired users to sophisticated content-based image retrieval and the scene understanding required by autonomous systems. Real-world photographs, the primary input for many computer vision tasks, typically depict complex scenes with multiple objects of varying sizes, categories, and semantic importance. These multi-object scenarios present a significant hurdle for <strong>vision-language models</strong>, which must accurately identify, localize (implicitly or explicitly), and then articulate all relevant entities in a coherent textual description. Datasets like Microsoft COCO <a href="https://arxiv.org/abs/1405.0312">Lin et al., 2014</a> are crucial benchmarks in computer vision as they exemplify this complexity by providing richly annotated, multi-object scenes. Understanding how advanced models such as CLIP <a href="https://arxiv.org/abs/2103.00020">Radford et al., 2021</a>—which learns joint embeddings from visual and textual data—handle these intricate visual interactions is therefore critical for improving the robustness, fairness, and overall performance of computer vision systems that aim to bridge the gap between pixels and semantics.</p>

<h2 id="literature-review">Literature Review</h2>

<h3 id="biases-in-multi-object-visionlanguage-models">Biases in Multi-Object Vision–Language Models</h3>

<p>Recent research has highlighted that vision–language models, especially CLIP (<a href="https://arxiv.org/abs/2103.00020">Radford et al., 2021</a>), exhibit notable biases when processing images containing multiple objects. These biases manifest primarily through the models’ tendencies to prioritize larger, visually dominant objects and objects mentioned earlier in captions (<a href="https://arxiv.org/abs/2502.19828">Abbasi et al., 2025</a>). Specifically, Abbasi et al. constructed controlled benchmarks (SimCO, CompCO) to analyze how varying object size and caption order affect CLIP’s image–text matching accuracy. Their results demonstrated that CLIP disproportionately focuses on larger objects visually, and textually prioritizes the first-mentioned object, significantly impacting multi-object captioning tasks.</p>

<p>Kamath et al. (2023) similarly explored CLIP’s limitations, revealing that its text encoder bottlenecks compositional information, causing failures in accurately representing object relationships, attribute–object bindings, and counts (<a href="https://arxiv.org/abs/2210.01936">Kamath et al., 2023</a>). For instance, CLIP struggled to differentiate semantically nuanced phrases like “a red square and a blue circle” from “a blue square and a red circle,” indicating weak compositional grounding.</p>

<p>Further studies on quantity biases reinforce these findings. Zhang et al. (2024) found that CLIP embeddings inadequately capture numerical object counts, leading to downstream errors such as incorrect object numbers in image generation (<a href="https://arxiv.org/abs/2310.01845">Zhang et al., 2024</a>). These biases collectively suggest that CLIP’s current embedding strategies offer limited fidelity for multi-object representation and grounding, emphasizing dominant objects while neglecting detailed compositional relationships.</p>

<h3 id="evaluation-methods-and-datasets">Evaluation Methods and Datasets</h3>

<p>To systematically study these biases, researchers have developed specialized datasets and analytical methods. <a href="https://huggingface.co/datasets/clip-oscope/simco-comco">Abbasi et al.’s (2025) SimCO and CompCO datasets</a> provide controlled synthetic and real-world image–caption pairs to precisely manipulate object prominence and mention order, making explicit the biases inherent in CLIP’s visual and textual encoders.</p>

<p>Other benchmarks like <strong>Winoground</strong> (Thrush et al., 2022) and <strong>CREPE</strong> (Yuksekgonul et al., 2023) specifically target relational and compositional understanding, exposing how small lexical or structural changes (e.g., swapping object positions in captions) significantly disrupt CLIP’s accuracy. Complementing these datasets, Chen et al. (2023) introduced <strong>gScoreCAM</strong>, a visualization technique that generates attention heatmaps highlighting CLIP’s focal regions within images (<a href="https://arxiv.org/abs/2302.05592">Chen et al., 2023</a>). Through such visualizations, researchers confirmed that CLIP frequently fixates on the most visually prominent object, further validating previous analytical findings and providing intuitive diagnostics for compositional failures.</p>

<h3 id="mitigation-techniques">Mitigation Techniques</h3>

<p>To address these compositional biases, researchers have proposed both architectural enhancements and novel training strategies. One notable approach is the integration of explicit object detection modules. For instance, <strong>MDETR</strong> (<a href="https://arxiv.org/abs/2104.12763">Kamath et al., 2021</a>) and <strong>GLIP</strong> (<a href="https://arxiv.org/abs/2112.03857">Li et al., 2022</a>) incorporate transformers trained to explicitly localize objects based on textual queries, thus promoting stronger grounding and compositional reasoning capabilities. These models substantially outperform traditional CLIP in tasks requiring precise multi-object correspondence.</p>

<p>Alternatively, Assouel et al. (2024) introduced <strong>OC-CLIP</strong>, an object-centric extension of CLIP designed specifically to enhance multi-object grounding. OC-CLIP utilizes slot-based object representations and graph-based scene parsing to achieve explicit bindings between image regions and caption components, significantly improving performance on compositional retrieval tasks (<a href="https://arxiv.org/abs/2311.11093">Assouel et al., 2024</a>).</p>

<p>Training-centric solutions have also been explored, including contrastive fine-tuning with carefully constructed hard negatives, and augmenting datasets with synthetic compositional examples. While effective in improving benchmark scores, these data-driven approaches address the symptoms of CLIP’s biases rather than the fundamental architectural limitations, underscoring the importance of combining architectural and data-centric solutions for robust multi-object vision–language modeling.</p>

<h3 id="our-take">Our take</h3>

<p>In the scope of a class term project, we’d like to probe the issues and ideate a potential post-hoc fix.</p>

<hr />

<h2 id="dataset---microsoft-coco">Dataset - Microsoft COCO</h2>

<p>For our analysis, we rely on the <a href="https://cocodataset.org/">Microsoft Common Objects in Context (COCO) dataset</a>—a widely used benchmark in computer vision research. COCO offers richly annotated images that capture complex everyday scenes, typically containing multiple objects of diverse categories, scales, and spatial relationships. This makes it well-suited for studying how vision-language models handle multi-object representations.</p>

<p>Each image in COCO includes multiple human-written captions and object-level annotations, including precise bounding boxes, object categories, and segmentation masks.</p>

<p>The dataset provides structured annotation fields such as:</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">image_id</code>: Unique identifier for each image</li>
  <li><code class="language-plaintext highlighter-rouge">bbox</code>: A bounding box in <code class="language-plaintext highlighter-rouge">[x, y, width, height]</code> format</li>
  <li><code class="language-plaintext highlighter-rouge">category_id</code>: Object class index (mapped to names via the category list)</li>
</ul>

<h3 id="defining-object-prominence">Defining Object Prominence</h3>

<p>In our study, we define the <em>prominence</em> of an object in an image as the ratio between the area of the object’s annotated bounding box and the total area of the image. This simple yet effective geometric measure quantifies an object’s spatial dominance within the visual scene. Prominence serves as a proxy for visual salience, under the assumption that larger objects are more likely to be visually or semantically prioritized by both human observers and vision-language models.</p>

<p>Formally, for an object with annotation <code class="language-plaintext highlighter-rouge">[x, y, width, height]</code> in the COCO dataset, and an image of dimensions <code class="language-plaintext highlighter-rouge">image_width</code> and <code class="language-plaintext highlighter-rouge">image_height</code>, prominence is calculated as:</p>

\[\begin{aligned}
\text{Prominence} = \frac{\text{width} \times \text{height}}{\text{image\_width} \times \text{image\_height}}
\end{aligned}\]

<p>Granted, an annotated bounding box area is not equivalent to the actual object’s size because not every object in an image appears in a perfect rectangle. While the COCO dataset does provide more granular segmentation annotations, we use the bounding box area as a simple proxy for the actual prominence of an object in a given image.</p>

<p>By quantifying prominence in this way, we can analyze whether models like CLIP are biased toward larger objects in multi-object scenes, especially when generating or scoring captions that describe such images.</p>

<h3 id="coco-categories">COCO Categories</h3>

<p><img src="/images/2025-05-13-6.8300-Final-Project-Blog/COCO_Categories.png" alt="COCO Categories Visualized" /></p>

<p>There are 80 annotated categories in the COCO <code class="language-plaintext highlighter-rouge">train2017</code> split. The diagram illustrates the hierarchical grouping of these object classes (shown in green) under broader super-categories (shown in blue), such as <code class="language-plaintext highlighter-rouge">animal</code>, <code class="language-plaintext highlighter-rouge">vehicle</code>, <code class="language-plaintext highlighter-rouge">kitchen</code>, and <code class="language-plaintext highlighter-rouge">electronic</code>. Each edge connects an object category to its corresponding super-category as defined by the COCO dataset’s metadata. This visualization highlights the diversity and complexity of the dataset.</p>

<h3 id="data-selection">Data Selection</h3>

<p>In this project, we selected <strong>502 images</strong> from the <strong>2017 training set</strong> (<code class="language-plaintext highlighter-rouge">train2017</code>).</p>

<p>Although the COCO <code class="language-plaintext highlighter-rouge">train2017</code> split contains 118K images, we had to select a subset based on the following criteria:</p>

<ol>
  <li><strong>Sufficient annotated objects per image:</strong> Our study focuses on multi-object scenarios, so images require a minimum number of objects. We required each image to have <strong>at least 3 annotated objects</strong>.</li>
  <li><strong>Sufficient prominence per object:</strong> After a brief exploration of the COCO dataset, we noticed some annotated objects can be very small. This rule aims to exclude instances where, for example, a tiny corner of a microwave is annotated as “microwave.” We required each annotated object to have <strong>at least 10% prominence</strong>.</li>
  <li><strong>Unique object categories per image:</strong> For example, constructing captions for images with multiple instances of the same object category (e.g., two “dogs”) would require disambiguation, which was outside the scope of this study. Therefore, we required all annotated objects in an image to belong to unique categories.</li>
</ol>

<table>
  <thead>
    <tr>
      <th>Filter Criterion</th>
      <th>Cumulative Images Remaining</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Initial <code class="language-plaintext highlighter-rouge">train2017</code> set</td>
      <td>118K</td>
    </tr>
    <tr>
      <td>&gt;= 3 objects per image</td>
      <td>81,982</td>
    </tr>
    <tr>
      <td>Each object &gt;= 10% prominence</td>
      <td>1,012</td>
    </tr>
    <tr>
      <td>All objects in unique categories</td>
      <td>502</td>
    </tr>
  </tbody>
</table>

<h2 id="captioning-method">Captioning Method</h2>

<p>For each of the 502 selected images, we employed the following captioning process:</p>

<ol>
  <li>Identify the <strong>top 3</strong> most prominent objects. E.g., <code class="language-plaintext highlighter-rouge">['airplane', 'truck', 'person']</code>. The order of these objects was varied in subsequent experiments.</li>
  <li>Generate a text caption by prefixing each object name with the appropriate English indefinite article (“a” or “an”) and conjoining them with “and”. E.g., <code class="language-plaintext highlighter-rouge">"an airplane and a truck and a person"</code>.</li>
</ol>

<h2 id="experiments">Experiments</h2>

<p>Inspired by the controlled multi-object bias study of (Abbasi et al., 2025), we evaluated whether the same prominence and order biases emerge when testing CLIP on <em>real-world</em>, high-resolution images from the COCO <code class="language-plaintext highlighter-rouge">train2017</code> dataset rather than on synthetic benchmarks like SimCO and CompCO. Whereas Abbasi et al. precisely manipulated object size and mention order in crafted scenes, we selected 502 COCO images containing at least three objects (each occupying ≥10% of the frame) and generated paired captions that vary the order of the top-3 objects or replace one. This real-data pipeline—defining “prominence” as the ratio of bounding-box area to image area and introducing a “random third object” for incorrect caption variants—allows us to probe CLIP’s biases in complex, natural settings using a simpler, accuracy-based metric.</p>

<p>Our results show a drop in matching accuracy from 91.23% (largest-first correct vs. random-third incorrect) to 87.45% when the correct caption is presented in smallest-first order, closely mirroring the performance degradation that (Abbasi et al., 2025) observed when swapping object mention order in synthetic scenes. Although the absolute magnitude of the drop is slightly attenuated—likely due to COCO’s richer visual context—this concordance suggests that CLIP’s order bias persists beyond controlled datasets and into natural, multi-object environments.</p>

<h3 id="1-correct-vs-incorrect-captions-with-largest-object-first">1. Correct vs. Incorrect Captions with Largest Object First</h3>

<p>We studied whether a less prominent object being misrepresented in a caption could mislead CLIP when scoring captions for the same image.</p>

<p>For each image, we constructed two captions:</p>

<ul>
  <li>A <strong>correct caption</strong> that mentions the top 3 most prominent objects in descending order of prominence (largest first).</li>
  <li>An <strong>incorrect caption</strong> that mentions the two most prominent objects correctly (in descending order of prominence) but replaces the third object with a randomly chosen category (ensuring it is unique from the other two) from the COCO dataset.</li>
</ul>

<p><img src="/images/2025-05-13-6.8300-Final-Project-Blog/000000150410.jpg" alt="COCO image 150410" /></p>

<p>For example, image <code class="language-plaintext highlighter-rouge">150410</code> (shown above) has the following three most prominent objects:</p>

<table>
  <thead>
    <tr>
      <th>Object</th>
      <th>Prominence</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">airplane</code></td>
      <td><code class="language-plaintext highlighter-rouge">0.252096262075527</code></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">truck</code></td>
      <td><code class="language-plaintext highlighter-rouge">0.2164387212748829</code></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">person</code></td>
      <td><code class="language-plaintext highlighter-rouge">0.10906725373243556</code></td>
    </tr>
  </tbody>
</table>

<p>We produced the following two captions:</p>

<ul>
  <li><strong>Correct</strong>: <em>“an airplane and a truck and a person”</em></li>
  <li><strong>Incorrect</strong>: <em>“an airplane and a truck and <strong>a chair</strong>“</em></li>
</ul>

<p>With all 502 images captioned with a pair of correct and incorrect captions, we asked CLIP to score each image against its two captions.</p>

<h4 id="result">Result</h4>

<p>CLIP preferred the correct caption for 458 out of 502 images—a 91.23% accuracy.</p>

<p>Here is an example failure:</p>

<p><img src="/images/2025-05-13-6.8300-Final-Project-Blog/000000069668.jpg" alt="COCO image 69668" /></p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>Caption</th>
      <th>CLIP Assigned Probability</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Correct</td>
      <td><em>“an oven and a person and a microwave”</em></td>
      <td>27.20%</td>
    </tr>
    <tr>
      <td>Incorrect</td>
      <td><em>“an oven and a person and <strong>a snowboard</strong>“</em></td>
      <td>72.80%</td>
    </tr>
  </tbody>
</table>

<p>This experiment suggests that while CLIP demonstrates a strong ability to identify correct multi-object captions (achieving 91.23% accuracy), its performance can be notably affected by inaccuracies related to less prominent objects. Instances where CLIP preferred an incorrect caption (where only the third most prominent object was altered) suggest that the model may not consistently ground or verify all listed objects with equal rigor. This implies that CLIP’s scoring mechanism may assign greater weight to dominant objects or that its compositional understanding is less robust for objects lower in the visual hierarchy.</p>

<h3 id="2-reordered-correct-vs-incorrect-captions">2. Reordered Correct vs. Incorrect Captions</h3>

<p>Building on the previous experiment, we studied whether CLIP’s caption scoring accuracy would further decrease if the correct caption listed objects from smallest to largest prominence.</p>

<p>For each image, we still constructed two captions:</p>

<ul>
  <li>A <strong>correct caption</strong> that mentions the top 3 most prominent objects in ascending order of prominence (smallest first).</li>
  <li>An <strong>incorrect caption</strong>: The incorrect caption (with the randomly swapped third object) from the previous experiment was reused.</li>
</ul>

<p>Still using the aforementioned image <code class="language-plaintext highlighter-rouge">150410</code> for example, the two captions were:</p>

<ul>
  <li><strong>Correct</strong>: <em>“a person and a truck and an airplane”</em> (objects are mentioned in reverse order of prominence compared to the previous experiment’s correct caption)</li>
  <li><strong>Incorrect</strong>: <em>“an airplane and a truck and <strong>a chair</strong>“</em></li>
</ul>

<h4 id="result-1">Result</h4>

<p>CLIP preferred the correct caption for 439 out of 502 images—an accuracy of 87.45%, lower than the previous experiment.</p>

<p>Although the absolute number of correctly-captioned images dropped by 19 (from 458 in the first experiment to 439 in this one), the change in performance is more intricate than a simple drop:</p>

<ul>
  <li>5 images whose <strong>incorrect captions</strong> were preferred by CLIP in the previous experiment: CLIP now preferred the correct caption.</li>
  <li>24 images whose correct captions were preferred by CLIP in the previous experiment: CLIP now preferred the <strong>incorrect caption</strong>.</li>
</ul>

<p>Image <code class="language-plaintext highlighter-rouge">273083</code> exemplifies one of these 24 cases where performance worsened:</p>

<p><img src="/images/2025-05-13-6.8300-Final-Project-Blog/000000273083.jpg" alt="COCO image 273083" /></p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>Caption</th>
      <th>CLIP Assigned Probability</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td> </td>
      <td>Previous experiment (Largest First Correct)</td>
      <td> </td>
    </tr>
    <tr>
      <td>Correct</td>
      <td><em>“a pizza and a dog and a dining table”</em></td>
      <td><strong>53.12%</strong></td>
    </tr>
    <tr>
      <td>Incorrect</td>
      <td><em>“a pizza and a dog and <strong>a cell phone</strong>“</em></td>
      <td>46.88%</td>
    </tr>
    <tr>
      <td> </td>
      <td>This experiment (Smallest First Correct)</td>
      <td> </td>
    </tr>
    <tr>
      <td>Correct</td>
      <td><em>“a dining table and a dog and a pizza”</em></td>
      <td>23.66%</td>
    </tr>
    <tr>
      <td>Incorrect</td>
      <td><em>“a pizza and a dog and <strong>a cell phone</strong>“</em></td>
      <td><strong>76.37%</strong></td>
    </tr>
  </tbody>
</table>

<p>This experiment reveals that the order in which objects are mentioned in a caption significantly impacts CLIP’s scoring, especially when less prominent objects are listed first. The accuracy decrease from 91.23% to 87.45% suggests CLIP is less confident or accurate when a correct caption’s textual sequence inverts the visual prominence hierarchy (i.e., smallest object mentioned first).</p>

<p>The misclassification of 24 images (previously scored correctly) when the correct caption was reordered indicates that CLIP might rely on an alignment between early-mentioned objects in the text and the most visually dominant objects in the image. When this alignment is disrupted (as in “smallest-first” correct captions), even if all objects are factually present, CLIP is more likely to prefer an incorrect caption that begins by mentioning the most prominent objects, even if it contains a subsequent error. This highlights a potential vulnerability: CLIP’s image-text alignment may be disproportionately influenced by a caption’s initial elements, potentially overshadowing a complete assessment of all described objects, especially when textual order mismatches visual salience.</p>

<h2 id="mitigation-caption-permutation-ensemble">Mitigation: Caption Permutation Ensemble</h2>

<p>To reduce CLIP’s sensitivity to the order in which objects are mentioned, we propose ensembling scores from all permutations of the 3 objects feature in a caption rather than relying on a single caption ordering. For an image \(I\) with top-3 objects \({o_1, o_2, o_3}\), the process is:</p>

<ol>
  <li>Form the set of all \(M=3! = 6\) permutations \(P = {\pi_1,\ldots,\pi_M}\), where each is a tuple of three objects.</li>
  <li>For each permutation \(\pi_m\), generate a caption \(c_m\) same as in previous experiments.</li>
  <li>
    <p>Compute its text embedding and normalize:
\(\begin{aligned}
tm \;=\;\mathrm{CLIP}_{\mathrm{text}}(c_m),\quad
\tilde t_m \;=\;\frac{t_m}{\|t_m\|_2}.
\end{aligned}\)</p>
  </li>
  <li>Average these \(M\) normalized embeddings to produce an order-invariant text representation:
\(\begin{aligned}
\bar t \;=\;\frac{1}{M}\sum_{m=1}^{M}\tilde t_m,
\qquad
\hat t \;=\;\frac{\bar t}{\|\bar t\|_2}.
\end{aligned}\)</li>
  <li>Encode and normalize the image once:
\(\begin{aligned}
v \;=\;\mathrm{CLIP}_{\mathrm{img}}(I),\quad
\hat v \;=\;\frac{v}{\|v\|_2}.
\end{aligned}\)</li>
  <li>Compute the final similarity score as the cosine similarity between \(\hat v\) and \(\hat t\):
\(\begin{aligned}
s_{\mathrm{ens}} \;=\;\langle \hat v,\;\hat t\rangle.
\end{aligned}\)</li>
</ol>

<p>We want to underscore that permutation ensemble method is applied to both the correct set of objects and the incorrect set of objects, hence making this a generalized method to use while in the real world, we don’t know if a provided list of objects are all factually featured in a given image.</p>

<p>By comparing similarity score for the correct set of objects to a similarly computed ensemble score for an incorrect set of objects (where the third object is consistently swapped across its permutations, as in previous experiments), we expect this ensemble approach to smooth out ordering noise and yield a more robust decision boundary.</p>

<h3 id="result-2">Result</h3>

<table>
  <thead>
    <tr>
      <th>Experiment</th>
      <th>Accuracy</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Experiment 1 (without mitigation)</td>
      <td>458 out of 502 (91.23%)</td>
    </tr>
    <tr>
      <td>With mitigation</td>
      <td><strong>452 out of 502 (90.04%)</strong></td>
    </tr>
    <tr>
      <td>Experiment 2 (without mitigation)</td>
      <td>439 out of 502 (87.45%)</td>
    </tr>
  </tbody>
</table>

<p>The permutation ensemble mitigation yields 90.04% accuracy (452/502), which sits between our original largest-first baseline (91.23%) and the reversed-order worst-case (87.45%). Although ensembling all six correct-caption permutations does not quite match the peak performance of the best single ordering, it substantially mitigates the drop seen with adverse orderings, recovering 2.59 percentage points (from 87.45% to 90.04%) compared to the smallest-first scenario. This demonstrates that averaging over permutations effectively smooths CLIP’s sensitivity to object mention order, trading a small amount of peak accuracy for markedly improved robustness against ordering noise.</p>

<hr />

<h2 id="conclusion">Conclusion</h2>

<p>Our investigation into CLIP’s handling of multi-object scenes using real-world COCO images confirms that the model, while generally proficient, exhibits clear sensitivities to both the accuracy of object mentions and their textual order. We found that:</p>

<ol>
  <li><strong>CLIP is vulnerable to inaccuracies concerning less prominent objects.</strong> While achieving a high accuracy (91.23%) when the most prominent objects were correctly listed first, errors in identifying the third object could still mislead the model, suggesting a potential hierarchy in how CLIP grounds objects.</li>
  <li><strong>Object mention order significantly impacts CLIP’s judgment.</strong> Reversing the caption order to list the smallest prominent object first (Experiment 2) led to a notable drop in accuracy to 87.45%. This demonstrates an “order bias,” where CLIP appears to favor captions that align with a “largest-first” heuristic, even if an alternative ordering is equally correct.</li>
  <li><strong>Permutation ensembling offers a viable mitigation strategy.</strong> By averaging the embeddings of all possible permutations of a correct object set, we achieved an accuracy of 90.04%. This approach successfully smoothed out the negative impact of unfavorable orderings, substantially recovering performance from the worst-case scenario (87.45%) and providing a more robust, order-agnostic representation. While this came at the cost of a slight dip from the optimal single-order accuracy, the gain in robustness is significant.</li>
</ol>

<p>These findings underscore the importance of considering object prominence and mention order when evaluating or deploying vision-language models like CLIP in complex, multi-object environments. While CLIP’s zero-shot capabilities are powerful, its internal biases can lead to performance variations that might be critical in real-world applications. The permutation ensemble method offers a practical step towards more reliable multi-object caption scoring, trading a small amount of peak performance for improved consistency.</p>

<p>Future work could explore a much bigger dataset size and more sophisticated ensembling techniques, investigate the architectural underpinnings of these biases within CLIP, or develop training strategies that inherently reduce such sensitivities, paving the way for even more robust and equitable vision-language understanding.</p>

<hr />

<h2 id="acknowledgement">Acknowledgement</h2>

<p>We want to <strong>credit</strong> and thank <a href="https://web.mit.edu/phillipi/">Professor Phillip Isola</a> for <strong>suggesting the permutation ensembling post-hoc approach</strong> during an insightful discussion about the project. It was truly an enlightening moment!</p>

<p>Special thanks to the <a href="https://www.scenerepresentations.org/courses/2025/spring/advances-in-cv/">Spring 2025 6.8300 teaching team</a>:</p>
<ul>
  <li><a href="https://www.vincentsitzmann.com/">Professor Vincent Sitzmann</a> for redesigning the course to be more relevant and for consistently bringing amazing energy to each lecture.</li>
  <li>Our wonderful TAs, <a href="https://vivekg.dev/">Vivek Gopalakrishnan</a>, <a href="https://yukaryote.github.io/">Isabella Yu</a>, and <a href="http://www.a14z.blog/">Adriano Hernandez</a>, providing invaluable feedback on this project and being open to engaging conversations about academic life.</li>
</ul>]]></content><author><name>Junru Ren 任俊儒</name><email>junru@computer.org</email></author><category term="class" /><category term="projects" /><category term="computer vision" /><category term="mit" /><summary type="html"><![CDATA[Abstract: Vision-language models like CLIP struggle with multi-object scenes, often favoring prominent objects or those mentioned first in captions. Using real-world COCO images, we show that CLIP’s caption-matching accuracy drops from 91.23% to 87.45% when object order is reversed. To address this, we explore a post-hoc mitigation: a permutation ensemble that averages scores across all object orders, boosting robustness and recovering accuracy to 90.04%. Our findings reveal persistent order biases and offer a simple, effective strategy to improve CLIP’s reliability in complex scenes.]]></summary></entry></feed>