Skip to main content

Command Palette

Search for a command to run...

#week 4 — Taming Logs, ASTs, and Graphs

Published
•4 min read•View as Markdown
#week 4  —  Taming Logs, ASTs, and Graphs
A
I’m a curious learner and aspiring software developer who enjoys building real-world systems, exploring new technologies, and learning in public. I write about coding, personal growth, and turning ideas into practical systems—while simplifying complex concepts for others.

This week in CodeAtlas was all about building the repo_parser, normalizing ASTs, wrangling Neo4j graphs, and surviving Node.js quirks. Lots of learning, debugging loops, and small victories!


1. Repo Parser: Building the Core Route

I started implementing the main API route for the repo_parser service:

Post  http://localhost:5001/generate-ast
body: {
    "repoName": "repo_name"
}

What it does:

  • Fetches filtered repository data from last week’s GitHub API work

  • Processes files for AST parsing

  • Prepares data for graph building

Challenges:

  • Node.js + nodemon restarted the server whenever logs were written

  • Logs for each file processed caused mid-processing restarts

  • Large repos → thousands of logs → hit GitHub API limits → endless restarts 😅

Solution:

Configured nodemon to ignore certain folders and log files:

{
  "ignore": [
    "public/",
    "public/**",
    "ast_test/",
    "*.log"
  ]
}

✅ Result: Server now runs smoothly without restarting on log writes


2. Node.js Event Loop & Parallel Processing

Once the repo parser was running, I experimented with parallel processing using Node’s event loop.

Setup:

  • Main service processing repo data

  • Log watcher service running alongside

Problem:

  • Log threshold reached while main service was busy

  • Event loop triggered server restart → endless loop

  • Functions stopped mid-execution → chaos

Lesson:

  • Node doesn’t have multiple threads; concurrency requires careful orchestration

  • Ignoring irrelevant files in nodemon prevents unnecessary restarts

  • Event loop management is key for processing large datasets


3. AST Normalization for Multiple Languages

Next focus: turning raw code files into normalized ASTs for graph building.

Key steps:

  • Single AST normalization didn’t scale → created language-specific normalization

  • Focused this week on JavaScript and TypeScript

  • Handled imports, exports, functions, and variables

Example snippet from a TypeScript file:

{
  "file": {
    "path": "string",           // Full path of the file
    "language": "string",       // Programming language
    "moduleType": "string",     // e.g., "module" or "script"
    "entryPoint": "boolean"     // Is this the entry point of the project?
  },
  "imports": [
    {
      "source": "string",       // Module/package being imported
      "kind": "string",         // "external" or "internal"
      "symbols": ["string"]     // Specific symbols/functions imported
    }
  ],
  "entities": {
    "variables": [
      {
        "name": "string",
        "kind": "string",       // "const", "let", "var"
        "valueType": "string"   // Type if known
      }
    ],
    "classes": [
      {
        "name": "string",
        "methods": ["string"],  // Method names in the class
        "properties": ["string"] // Properties in the class
      }
    ],
    "functions": [
      {
        "id": "string",         // Unique ID for the function (path + name)
        "name": "string",
        "scope": "string",      // "global" or "local"
        "params": ["string"],   // List of parameters
        "calls": ["string"]     // Functions or methods called inside
      }
    ],
    "modules": ["string"]       // Sub-modules or nested modules
  },
  "exports": [
    {
      "name": "string",          // Exported variable/function/class name
      "kind": "string",          // "variable", "function", "class"
      "default": "boolean"       // Is it a default export?
    }
  ]
}

Lessons Learned:

  • export default handling in TS/JS is subtle but crucial

  • Normalized ASTs allow consistent downstream graph building


4. Neo4j Graph Database Challenges

Building the codebase graph brought its own hurdles:

  • Neo4j Community Edition can’t create multiple databases

  • Had to reset the entire database for new projects

  • Restarting servers mid-processing caused loops if not handled carefully

Takeaways:

  • Plan DB resets and graph creation carefully

  • Combine with nodemon ignore rules to prevent server loops

  • Visualizing ASTs and graphs becomes much smoother once normalized


Key Takeaways

  • Configure nodemon to ignore log and build folders

  • AST normalization must be language-specific

  • Event-loop management is essential for processing large datasets in Node

  • Neo4j Community Edition has limitations — plan accordingly

  • GitHub API limits require batching & caching strategies


What’s Coming Next

Next week, my focus on.

Handling GitHub API Edge Cases

  • Implement a middleware to check if the user has enough API quota before forwarding requests.

  • Prevent unnecessary API calls if limits are exceeded.

  • Handle errors gracefully when GitHub responses fail or are incomplete.

Managing Repository Updates

  • Detect when a repository has new commits or file changes.

  • Update the stored content in MySQL and the AST/graph without reprocessing everything from scratch.

  • Ensure incremental updates are efficient and consistent.

Resuming Interrupted Processes

  • If a repo parsing process stops midway (e.g., server restart, crash, or threshold reached), automatically resume from the last processed file.

  • Keep track of progress and avoid duplicating work.

Frontend & Visualization

  • If there’s time remaining, I’ll start building the frontend to visualize repositories, ASTs, and dependency graphs.

  • Focus on creating interactive and easy-to-navigate views for developers.


🤝 Contributions

Ideas, improvements, and suggestions are always welcome.
You’re encouraged to submit issues or pull requests to help evolve the platform.

Github repo

CodeAtlas

Part 5 of 11

CodeAtlas is an AI-powered system that analyzes GitHub repos, generates AST-based insights, builds code relationship graphs, and helps developers understand complex projects through visualization and intelligent search.

Up next

#week 5 — Scaling, Stability & Production Reality Checks

This week in CodeAtlas was about moving from “it works” to “it survives real-world usage.” I focused on GitHub rate limits, backend graph exploration, logging, data integrity, and—most importantly—running a hard production readiness review that expos...