Skip to main content

Command Palette

Search for a command to run...

#week 3 — Diving Into GitHub APIs, Visualizing Logs in Grafana & Adding Repo File Processing

Published
•5 min read•View as Markdown
#week 3 — Diving Into GitHub APIs, Visualizing Logs in Grafana & Adding Repo File Processing
A
I’m a curious learner and aspiring software developer who enjoys building real-world systems, exploring new technologies, and learning in public. I write about coding, personal growth, and turning ideas into practical systems—while simplifying complex concepts for others.

This week was full of learning, confusion 😅, breakthroughs, and new features for CodeAtlas.
I continued building on last week’s logging system and started working with GitHub APIs for the first time — which was not easy at all.

Here’s everything I accomplished this week 👇


1️⃣ Understanding the GitHub API (The Hard Way)

This week I started working with the GitHub REST API using the official docs:

https://docs.github.com/en/rest?apiVersion=2022-11-28

But honestly… the docs confused me a lot at the beginning 😅
I couldn't understand half the endpoints or how authentication worked.

So I switched to watching YouTube tutorials, reading small examples, and using ChatGPT & Gemini for explanations.

Slowly, everything started making sense:

✔ How to authenticate
✔ How to fetch repository data
✔ How rate limits work
✔ How to structure requests properly

It was frustrating at first, but finally I was able to make successful requests and integrate them into CodeAtlas.


2️⃣ Connecting Logs to Grafana (Real-Time Visualization 🎉)

Last week, I built a full logging system → Winston → MySQL.
This week, I took it one step further:

✅ Logs are now visible in Grafana .

This means:

  • I can see logs from every microservice on a dashboard

  • Filtering and searching logs is super easy

  • No more digging through MySQL manually

  • Errors now pop up instantly

  • I can monitor the whole system visually

This alone made debugging and tracking behavior 10x easier.

User Setup Creating the dedicated grafana user account in MySQL Workbench.

User Setup Creating the dedicated grafana user account in MySQL Workbench.

Access Control Granting read-only permissions (SELECT) to the log database to ensure security.

Access Control Granting read-only permissions (SELECT) to the log database to ensure security.

Data Visualization Visualizing log severity metrics in Grafana using the MySQL data source.

Data Visualization Visualizing log severity metrics in Grafana using the MySQL data source.


3️⃣ Fetching All Files From a GitHub Repo

Next, I worked on the second API in my system:

http://localhost:5000/repo?owner=<owner>&repo=<repo>

/repos/${owner}/${repo}/git/trees/${defaultBranch}?recursive=1

The goal:
Fetch all files from a GitHub repository → process them → store useful ones → send them to Kafka

While building this, I had to solve multiple problems:

✔ Step 1 — Fetch all repo data from GitHub

I used the GitHub Contents API to get the full file tree.

✔ Step 2 — Store results in MySQL

Before inserting anything, I added a custom filter.

✔ Step 3 — Added filtering to skip unwanted files

I avoided loading:

  • images (jpg, png, svg, etc.)

  • binary files

  • config and environment files

  • anything over a certain size

This keeps the dataset clean and reduces noise.


// --- File Filtering Logic ---
const IGNORED_EXTENSIONS = new Set([
    // Media (Images, Video, Audio)
    '.jpg', '.jpeg', '.png', '.gif', '.bmp', '.ico', '.svg', '.tiff', '.webp',
    '.mp4', '.mkv', '.avi', '.mov', '.wmv', '.flv', '.webm',
    '.mp3', '.wav', '.aac', '.flac', '.ogg', '.m4a',
    // Documents & Archives
    '.pdf', '.doc', '.docx', '.xls', '.xlsx', '.ppt', '.pptx',
    '.zip', '.rar', '.7z', '.tar', '.gz', '.bz2', '.iso',
    '.md', '.markdown', '.txt', '.rst', // Docs
    // Binaries & Bytecode
    '.exe', '.dll', '.so', '.dylib', '.bin', '.obj', '.o', '.a', '.lib',
    '.pyc', '.class', '.jar', '.war',
    // Logs & DB
    '.log', '.sqlite', '.db',
    // Font files
    '.ttf', '.otf', '.woff', '.woff2', '.eot',
    // Web Assets (often noise for code analysis)
    '.css', '.scss', '.less', '.html', '.htm', '.map'
]);

const IGNORED_FILES = new Set([
    'package-lock.json', 'yarn.lock', 'pnpm-lock.yaml', 'bun.lockb',
    'Cargo.lock', 'Gemfile.lock', 'composer.lock',
    '.DS_Store', 'Thumbs.db', '.env', '.env.local',
    'Dockerfile', 'docker-compose.yml', 'LICENSE', 'README.md',
    'Makefile', 'CMakeLists.txt' // Build files
]);

// Directories to skip entirely
const IGNORED_DIRS = new Set([
    'node_modules', 'bower_components', 'jspm_packages',
    'venv', '.venv', 'env',
    'dist', 'build', 'out', 'target', 'bin', 'obj',
    '.git', '.svn', '.hg', '.idea', '.vscode', '.settings', '.next', '.nuxt',
    'coverage', '__tests__', 'test', 'tests',
    'public', 'assets', 'static', 'resources', 'images', 'img', 'media', 'videos' // Asset folders
]);

Final Output Verification Querying the restoree table in MySQL Workbench to confirm the filtered repository structure and metadata have been successfully populated.

✔ Step 4 — Publish the filtered data to Kafka

After cleaning, only meaningful files (code files, markdown files, etc.) are:

  • sent to Kafka for background processing

  • saved for further analysis like dependency mapping, AST parsing, etc.

This is a big step for my upcoming “repository explorer” feature.


💡 What I Learned This Week

  • GitHub APIs are powerful but not beginner-friendly

  • Grafana dashboards make logs feel alive

  • Data filtering is crucial when dealing with thousands of repo files


🔮 What’s Coming Next

Next week, my main focus will be building one of the most important parts of CodeAtlas:
👉 the repo_parser core service.

This service will:

  • take filtered repository files

  • understand project structure

  • extract meaningful metadata

  • power future features like code analysis and navigation

But there’s a real-world constraint I hit this week:

⚠️ GitHub API rate limits.

While testing repository fetching at scale, I hit GitHub’s API limits — a reminder that real systems don’t run in ideal conditions.

So alongside repo_parser, I’ll also be working on:

  • smarter API usage

  • request batching & caching

  • better rate-limit handling and backoff strategies

Because before a system can be smart, it has to be reliable.

If CodeAtlas is going to scale, it needs to respect the limits of the platforms it depends on.

More lessons coming next week 🚀


🤝 Contributions

Ideas, improvements, and suggestions are always welcome.
You’re encouraged to submit issues or pull requests to help evolve the platform.

Github repo

CodeAtlas

Part 4 of 11

CodeAtlas is an AI-powered system that analyzes GitHub repos, generates AST-based insights, builds code relationship graphs, and helps developers understand complex projects through visualization and intelligent search.

Up next

#week 4 — Taming Logs, ASTs, and Graphs

This week in CodeAtlas was all about building the repo_parser, normalizing ASTs, wrangling Neo4j graphs, and surviving Node.js quirks. Lots of learning, debugging loops, and small victories! 1. Repo Parser: Building the Core Route I started implemen...