r/MachineLearning 2d ago

Research Claude Code for Research Papers [R]

Third-year PhD student, NLP / interpretability. I want a reality check from people doing similar work.

I started using Claude Code for the boring parts: argparse boilerplate, plotting, config wrangling. Over the last few months the scope has crept. It now writes most of my experiment scaffolding, refactors my dataloaders, does first-pass debugging on training runs, and drafts the analysis scripts. I mostly read diffs and say yes.

The output is fine. My throughput is up. The thing bothering me is that I no longer hold my own codebase in my head. When a result looks off, I used to have an instinct about which line was lying to me. Now I go hunting like it’s someone else’s repo. I catch bugs later than I used to, and I catch them by reasoning about the numbers rather than by knowing the code.

I don’t think the tool is the problem. I think I delegated a layer that was doing more for my understanding than I gave it credit for.

Questions for people further along or in the same spot:

  1. Roughly what fraction of your research code do you write yourself now?

  2. Is there anything you deliberately refuse to hand off? (For me I think the eval harness and anything defining a metric should stay mine, but I keep breaking my own rule.)

  3. Does anyone have a workflow that keeps the speedup without the detachment? Reading the diff line by line is not cutting it.

Not looking for a “tools are just tools” answer. I’m asking about the specific feeling of not owning your own experiments anymore.

260 Upvotes

62 comments sorted by

View all comments

45

u/milesper 2d ago

Fifth year PhD and current research intern— I use it extensively to debug code and suggest ideas, but the code it actually writes is a very confined scope—data analysis and visualizations. For boilerplate like configs, I have a few files I reuse across projects that I can quickly modify for the new project.

19

u/milesper 2d ago

My feeling is that while agents can certainly write the code fine in most cases, my goal as a researcher is to understand the problem and method as well as possible, and one great way to do that is to actually implement everything. And frankly most ML experiments are pretty easy to code up nowadays so it’s not really a big productivity loss.