Project Info
In the spirit of innovation and agility, LogFlowAI was seeded in our mind in a meta sense during our internship and work experiences at AWS & TikTok respectively—a project that melds deep technical expertise with a sharp business acumen to revolutionize how organizations handle system logs and preempt failures. Below, we share our journey through inspiration, development, challenges, accomplishments, learnings, and our ambitious plans ahead. Our mission was clear: revolutionize the way organizations handle system logs by transforming raw data into actionable business insights.
Inspiration
During our time in the industry, we saw how reactive troubleshooting could cost valuable time and resources. Inspired by the potential of predictive analytics and real-time monitoring, we asked: What if we could forecast system failures and diagnose issues before they escalate? This vision led us to create a platform that marries deep technical prowess with strategic business insights—empowering teams to move from crisis management to proactive decision-making. What It Does LogFlowAI is an end-to-end solution designed to modernize log analysis and system monitoring by combining automation with state-of-the-art AI. Key capabilities include: Real-Time Log Ingestion & Processing: Seamlessly ingests log data from diverse sources—ranging from CSV files and syslog streams to JSON endpoints and Kubernetes logs—with support for both batch processing and sub-second real-time streaming. Real-Time Log Ingestion & Processing: Seamlessly ingests log data from diverse sources—ranging from CSV files and syslog streams to JSON endpoints and Kubernetes logs—with support for both batch processing and sub-second real-time streaming. Advanced Filtering & Analytics: Empowers teams to sift through millions of log entries using sophisticated search capabilities, including full-text search, regex patterns, and custom filters. This enables pinpointing issues quickly and accurately. Advanced Filtering & Analytics: Empowers teams to sift through millions of log entries using sophisticated search capabilities, including full-text search, regex patterns, and custom filters. This enables pinpointing issues quickly and accurately. Predictive Insights & Neural Network Predictions: Leverages customized advanced machine learning models (built manually with deep learning frameworks) trained on historical data and large-scale HDFS trace benchmarks to predict system anomalies and potential failures. Predictive Insights & Neural Network Predictions: Leverages customized advanced machine learning models (built manually with deep learning frameworks) trained on historical data and large-scale HDFS trace benchmarks to predict system anomalies and potential failures. Interactive Dashboard Visualizations: Converts complex log data into clear, actionable insights through dynamic dashboards. Visualizations include real-time metrics, heat maps, and time-series graphs that help teams monitor service health and error rates at a glance. Interactive Dashboard Visualizations: Converts complex log data into clear, actionable insights through dynamic dashboards. Visualizations include real-time metrics, heat maps, and time-series graphs that help teams monitor service health and error rates at a glance. Extensible API & Seamless Integration: Offers robust RESTful endpoints to integrate with existing infrastructure, enabling automated analyses and streamlined reporting across various business systems. Extensible API & Seamless Integration: Offers robust RESTful endpoints to integrate with existing infrastructure, enabling automated analyses and streamlined reporting across various business systems. This combination of capabilities not only enhances operational efficiency but also translates technical insights into strategic business intelligence for sales, operations, and executive teams. How We Built It Our approach was to design a modular, scalable architecture that unifies diverse technologies into one cohesive platform: Modular Architecture & Data Orchestration: We integrated robust data ingestion pipelines with a suite of management commands and endpoints, ensuring that log data—regardless of its origin—is processed efficiently and accurately. Modular Architecture & Data Orchestration: We integrated robust data ingestion pipelines with a suite of management commands and endpoints, ensuring that log data—regardless of its origin—is processed efficiently and accurately. Hybrid Data Processing: By supporting both batch and real-time streaming, LogFlowAI offers sub-second latency in log analysis while handling vast volumes of data, ensuring that teams always have the most current insights. Hybrid Data Processing: By supporting both batch and real-time streaming, LogFlowAI offers sub-second latency in log analysis while handling vast volumes of data, ensuring that teams always have the most current insights. AI-Powered Predictive Analytics: A neural network model, trained on extensive datasets including HDFS trace benchmarks, drives our predictive capabilities. This model empowers the system to forecast potential issues before they occur, enabling proactive maintenance. AI-Powered Predictive Analytics: A neural network model, trained on extensive datasets including HDFS trace benchmarks, drives our predictive capabilities. This model empowers the system to forecast potential issues before they occur, enabling proactive maintenance. Complex Computational Graph: Our Langchain computational graph features an orchestrator LLM that directs the overall workflow by delegating tasks to specialized modules. Three sub-LLMs, gaining information from our backend neural network, handle specific analyses, ensuring comprehensive coverage of our log data. Their outputs are then merged and refined by an orchestrator-combine LLM, delivering cohesive, actionable insights. Complex Computational Graph: Our Langchain computational graph features an orchestrator LLM that directs the overall workflow by delegating tasks to specialized modules. Three sub-LLMs, gaining information from our backend neural network, handle specific analyses, ensuring comprehensive coverage of our log data. Their outputs are then merged and refined by an orchestrator-combine LLM, delivering cohesive, actionable insights. Integrated Visual & API Layers: Our interactive dashboards and comprehensive API provide intuitive access to key metrics and allow seamless integration with other business tools, ensuring that insights are both accessible and actionable across the organization. Integrated Visual & API Layers: Our interactive dashboards and comprehensive API provide intuitive access to key metrics and allow seamless integration with other business tools, ensuring that insights are both accessible and actionable across the organization. Quality & Scalability Focus: We built our solution with continuous integration and deployment in mind, ensuring that as the system scales, it remains reliable, secure, and easy to maintain. Quality & Scalability Focus: We built our solution with continuous integration and deployment in mind, ensuring that as the system scales, it remains reliable, secure, and easy to maintain. Challenges We Ran Into Building a platform as comprehensive as LogFlowAI presented several challenges: Data Scale & Integrity: Processing massive log datasets while ensuring data quality and consistency demanded innovative parsing and validation techniques. Data Scale & Integrity: Processing massive log datasets while ensuring data quality and consistency demanded innovative parsing and validation techniques. Seamless Integration: Orchestrating various components—ranging from real-time ingestion and filtering to predictive analytics and visualization—required meticulous synchronization and robust architecture design. Seamless Integration: Orchestrating various components—ranging from real-time ingestion and filtering to predictive analytics and visualization—required meticulous synchronization and robust architecture design. Balancing Technical Innovation with Business Value: Translating complex technical insights into actionable, easy-to-understand business intelligence required iterative design and close collaboration with potential users to ensure relevance and usability. Balancing Technical Innovation with Business Value: Translating complex technical insights into actionable, easy-to-understand business intelligence required iterative design and close collaboration with potential users to ensure relevance and usability. Accomplishments We’re Proud Of Transformative Predictive Analytics: Successfully integrating AI to forecast system issues, thereby shifting the paradigm from reactive troubleshooting to proactive system management. Transformative Predictive Analytics: Successfully integrating AI to forecast system issues, thereby shifting the paradigm from reactive troubleshooting to proactive system management. Real-Time Operational Insights: Delivering dynamic dashboards and advanced filtering capabilities that empower teams to monitor and act on system performance instantly. Real-Time Operational Insights: Delivering dynamic dashboards and advanced filtering capabilities that empower teams to monitor and act on system performance instantly. Business-Driven Innovation: Bridging the gap between technical data and strategic decision-making, enabling non-technical stakeholders to make informed, forward-looking decisions. Business-Driven Innovation: Bridging the gap between technical data and strategic decision-making, enabling non-technical stakeholders to make informed, forward-looking decisions. Robust, Scalable Architecture: Building a solution that is not only technologically advanced but also engineered for reliability and scalability across diverse enterprise environments. Robust, Scalable Architecture: Building a solution that is not only technologically advanced but also engineered for reliability and scalability across diverse enterprise environments. What We Learned Our journey with LogFlowAI has underscored several key lessons: Integration is Essential: The seamless fusion of various data sources and analytic tools is critical for delivering accurate and actionable insights. Integration is Essential: The seamless fusion of various data sources and analytic tools is critical for delivering accurate and actionable insights. Automation Fuels Efficiency: Implementing automated data pipelines, testing, and deployment processes is indispensable for maintaining high-quality software and operational resilience. Automation Fuels Efficiency: Implementing automated data pipelines, testing, and deployment processes is indispensable for maintaining high-quality software and operational resilience. User-Centric Design is Paramount: Bridging the divide between complex technical processes and business strategy requires a clear focus on usability and actionable reporting. User-Centric Design is Paramount: Bridging the divide between complex technical processes and business strategy requires a clear focus on usability and actionable reporting. Innovative Thinking Drives Business Impact: Leveraging advanced AI and real-time analytics not only solves technical problems but also creates tangible business value by anticipating challenges and guiding strategic decisions. Innovative Thinking Drives Business Impact: Leveraging advanced AI and real-time analytics not only solves technical problems but also creates tangible business value by anticipating challenges and guiding strategic decisions. What’s Next for LogFlowAI Our journey is far from over. We envision a future where LogFlowAI continues to push the boundaries of what’s possible: Enhanced Predictive Analytics: Further refining our machine learning models to improve the accuracy of system failure predictions and anomaly detection. Enhanced Predictive Analytics: Further refining our machine learning models to improve the accuracy of system failure predictions and anomaly detection. Dynamic Real-Time Monitoring: Developing even more interactive and customizable dashboards to provide deeper insights into system performance and business impact. Dynamic Real-Time Monitoring: Developing even more interactive and customizable dashboards to provide deeper insights into system performance and business impact. Expanded Data Integration: Broadening our support for additional log sources and data streams to offer a more comprehensive view of operational health. Expanded Data Integration: Broadening our support for additional log sources and data streams to offer a more comprehensive view of operational health. Scalability & Enterprise Readiness: Optimizing performance and security to ensure our platform can scale effortlessly for larger organizations with complex infrastructures. Scalability & Enterprise Readiness: Optimizing performance and security to ensure our platform can scale effortlessly for larger organizations with complex infrastructures. Deepened AI Insights for Business Strategy: Enhancing our generative AI capabilities to deliver even more nuanced, business-focused insights that drive strategic planning and growth. Deepened AI Insights for Business Strategy: Enhancing our generative AI capabilities to deliver even more nuanced, business-focused insights that drive strategic planning and growth. Closing Thoughts LogFlowAI is more than a log analysis tool—it’s a vision for a future where data drives every decision and every log tells a story. By combining real-time analytics, advanced AI, and user-centric design, we are transforming the way organizations approach DevOps and business strategy. As we continue to innovate and expand, our commitment remains to empower teams with the foresight and agility needed to thrive in an ever-evolving digital landscape. Let's build a future where proactive insights replace reactive firefighting, and where every byte of data fuels smarter, strategic decisions.
LogFlowAI - Real-time Log Analysis & Monitoring
LogFlowAI is a web application created for real-time log analysis and monitoring. The platform ingests, filters, and visualizes log data to help teams diagnose issues and monitor system performance effectively. This project also demonstrates advanced capabilities ranging from deep learning using PyTorch to hybrid data operations via an InterSystems IRIS vector store, alongside large-scale log/data analytics using HDFS trace benchmarks.
Key Features
- Real-time Log Ingestion: Stream and process logs from diverse sources including CSV files, syslog streams, JSON endpoints, and Kubernetes logs. Support both batch processing for historical data and real-time streaming with sub-second latency. Automatically handle data validation, field extraction, and timestamp normalization across multiple timezones.
- Advanced Filtering: Power through millions of log entries with our sophisticated search capabilities. Combine full-text search, regex patterns, and field-specific filters to pinpoint exactly what you need. Features include fuzzy matching, saved searches, query templates, and support for complex boolean operations across log level, service name, timestamp, and custom fields.
- Dashboard Visualizations: Transform raw logs into actionable insights with our interactive dashboards. Track key metrics like error rates and service health in real-time, visualize log patterns with heat maps and time-series graphs, and create custom views for different teams. All visualizations support drill-down capabilities and export options for deeper analysis.
- Extensible API: Seamlessly integrate LogFlowAI into your existing infrastructure through our comprehensive RESTful API. Ingest logs programmatically, trigger automated analyses, and export metrics to external systems. The API includes robust authentication, rate limiting, detailed documentation, and support for custom plugins to extend functionality.
- Neural Network Architecture: A deep learning model built in PyTorch, trained on CSV files (including preprocessed HDFS trace benchmarks) to predict outcomes based on historical data.
- Database Architecture: An IRIS vector store is employed for advanced document storage and similarity searches using SQL-backed operations.
- HDFS Trace Bench Data: Large-scale HDFS trace benchmarks provide performance metrics and error patterns, offering rich data for analytics and model training.
Getting Started
To clone and set up the application locally, follow the instructions below.
Clone the Repository
git clone https://github.com/DhyeyMavani2003/logflowai.git
cd logflowai
Set Up a Virtual Environment
Create and activate a virtual environment to manage dependencies locally.
python3 -m venv env
source env/bin/activate
Install Dependencies
pip install -r requirements.txt
Run Database Migrations
Apply migrations to set up the database schema.
python manage.py makemigrations
python manage.py migrate
Import Log Data
Your log data (e.g., CSV files) is located under logapp/data/. You can import logs using the built-in management command. For example, to import logs from a CSV:
python manage.py import_logs
Or, you can trigger the import from the LogFlowAI home page using the "Import Logs from CSV" button.
Start the Development Server
To run the application locally, use the following command:
python manage.py runserver
Visit http://127.0.0.1:8000/ in your browser to view the application.
System Architecture & Design
LogFlowAI is built with a modular and scalable architecture. Below is a high-level diagram of the core components and their interactions.

Architecture Components
- LogDB: The central database storing log entries for efficient querying.
- Importer: A suite of management commands and endpoints to ingest log data from various sources.
- Home: Displays log entries with advanced filtering options.
- Dashboard: Provides aggregated visual insights (e.g., logs per hour, unique services).
- API: RESTful endpoints to trigger log imports and fetch filtered log data.
Functional Design
The following diagram provides a detailed breakdown of the application’s internal components and workflows.
graph TD
%% Data Ingestion Layer
subgraph "Data Ingestion"
A1[Various Log Sources CSV, Syslog, JSON, Kubernetes]
A2[Log Importer Management Commands & Endpoints]
A1 --> A2
end
%% Data Storage Layer
subgraph "Data Storage"
B1["LogDB PostgreSQL/SQLite"]
B2["IRIS Vector Store Document Embeddings"]
end
%% Backend Processing Layer
subgraph "Backend Processing"
C1[Log Filtering & Aggregation]
C2[Dashboard Metrics Computation]
C1 --> B1
C2 --> B1
end
%% Machine Learning & Predictions
subgraph "Machine Learning"
M1[Neural Network PyTorch]
M2[Training Data:<br>CSV & HDFS Trace Benchmarks]
M2 --> M1
M1 --> B1
end
%% Document Ingestion & Similarity Search
subgraph "Search & Similarity"
S1[Document Ingestion<br>& Embedding]
S2[Similarity Search]
S1 --> B2
S2 --> B2
end
%% API & Integration Layer
subgraph "API & Integration"
A3[RESTful API & FastAPI<br>/predict Endpoint]
A3 --> C1
A3 --> C2
A3 --> M1
A3 --> S2
end
%% Frontend Layer
subgraph "Frontend"
F1[Home Page:<br>Log Listing & Filtering]
F2[Dashboard:<br>Visualizations & Metrics]
F1 --> C1
F2 --> C2
end
Neural Network Architecture
The model is implemented in PyTorch, using data loaded from several CSV files including standard splits (train, val, test) and benchmark data. A custom CSVDataset processes numerical features and converts targets into tensors.
Dataset and Preprocessing
- Data is sourced from multiple CSV files:
- Standard splits: train.csv, val.csv, test.csv
- Benchmark dataset: rowNumberResult.csv (contains performance metrics from HDFS traces)
- Preprocessing involves cleaning data, extracting numerical features, and preparing tensors for model input.
Model Architecture
The neural network (encapsulated in the Net class) includes:
- Input Layer: Accepts a dynamic set of features.
- Four Hidden Layers:
- Layer 1: Linear (input → 128) → BatchNorm → ReLU → Dropout (0.3)
- Layer 2: Linear (128 → 64) → BatchNorm → ReLU → Dropout (0.3)
- Layer 3: Linear (64 → 32) → BatchNorm → ReLU → Dropout (0.3)
- Layer 4: Linear (32 → 16) → BatchNorm → ReLU → Dropout (0.3)
- Output Layer: A final linear layer producing a single output, transformed via sigmoid activation to yield probabilities.
Training and Evaluation
- Loss Function: Utilizes nn.BCEWithLogitsLoss() to handle raw logits.
- Optimizer: Adam optimizer with a learning rate typically set to 0.001.
- Metrics: Tracks training loss, accuracy, True Positive Rate (TPR), and True Negative Rate (TNR).
- Model Saving: The trained weights are saved as model.pth for deployment and inference.
Deployment with FastAPI
- A FastAPI server (app.py) is set up to load the trained model and serve predictions via a /predict endpoint.
- The endpoint expects JSON input (a list of feature values) and returns the predicted probability and class.
Database and Vector Store Architecture
This project leverages an advanced vector store implemented with the InterSystems IRIS driver:
- Vectorized Document Storage: Documents are embedded and stored in a persistent SQL table.
- Robust Search and Reconnection: The IRIS-based retriever supports similarity searches, ensuring persistent connections to the vector store.
IRIS Vector Store Overview
Documents are transformed into embeddings and stored within the IRIS vector store. This allows for efficient similarity searches and analytical queries using SQL-based operations.
Document Ingestion and Similarity Search
- Ingestion: Demonstrated in notebooks like langchain_demo.ipynb, where documents are processed and embedded.
- Search: The embedded vectors are used to perform similarity searches, facilitating quick retrieval of related documents.
HDFS Trace Bench Data
The project incorporates HDFS trace benchmarks, particularly through files like rowNumberResult.csv, which contains measurements and performance metrics. These benchmarks are used for:
- Analyzing performance characteristics.
- Feeding relevant metrics into the neural network.
- Enriching the document store for similarity queries.
Quickstart
-
Clone the Repository:
bash git clone https://github.com/yourusername/treehacks-2025.git cd treehacks-2025 -
Set Up the Environment:
bash python -m venv venv source venv/bin/activate # On Windows: venv\Scripts\activate -
Install Dependencies:
bash pip install -r requirements.txt -
Run Data Preprocessing and Training:
bash python preprocess.py python train.py -
Start the FastAPI Server:
bash uvicorn app:app --reload -
Open and Run the Notebooks: Explore Jupyter notebooks such as: - main.ipynb for training and evaluation. - langchain_demo.ipynb for document ingestion and similarity search features.
Contributing
We welcome contributions to improve LogFlowAI. Please follow these guidelines:
-
Code Formatting:
Ensure your code adheres to the Black style guidelines. Format your code by running:python -m black . -
Documentation:
For adding or updating documentation, refer to the LogFlowAI Documentation Guide in thedocs/directory. This includes instructions on using Sphinx for documenting models and views. -
Pull Requests:
Before submitting a pull request, ensure your code is well-documented, tested, and follows our coding standards.
For any questions or further support, please reach out to our team. Enjoy using LogFlowAI for your log analysis needs!
Analysis
View
Metric
- 35
- 8
- 2
Figures cover GitHub contributors during the hackathon window. A co-authored commit counts in full for each author, so per-member totals add up to more than the whole-team figures.
Technology
- AnthropicIn code
- CSSIn code
- DjangoIn code
- FastAPIIn code
- HTMLIn code
- LangChainIn code
- OpenAIIn code
- PythonIn code
- PyTorchIn code
- DockerClaimed
- JavaScriptClaimed
9 of 11 appear in the indexed code. 2 claimed on Devpost could not be matched to code, which may simply mean the tool leaves no trace in the repository.
AI coding agents
No AI coding agent signals were found in this repository.
Detected from committed agent config files and commit authorship. Absence of a signal is not proof an agent was unused.
Codebase size
Source size
109 KB
Source files
52
Counts recognized source files only; vendored directories, binaries and lockfiles are excluded, so this is smaller than the repository on disk.
Repository
DhyeyMavani2003/logflowai
75 files · 23.0 MB · @ 792deb1
Structure
Interface
2 files · 3%Screens, components and styles rendered to the user.
API & routing
1 file · 1%Request entry points: routes, handlers and controllers.
Application logic
28 files · 37%Domain rules, services and shared utilities.
Data & schema
6 files · 8%Schema definitions, migrations and data access.
Supporting
Layers are inferred from where files sit in the tree, not from reading the code. A project that names its directories unconventionally will read oddly here — open the file browser to check anything the diagram implies.
Languages
- Python49%
- Markdown39%
- HTML7%
- YAML5%
- CSS0%
Share of indexed source by file size. Binary and vendored files are excluded.
Dependencies
requirements.txt
pypi · 128- aiohappyeyeballs
- aiohttp
- aiosignal
- alabaster
- amqp
- annotated-types
- anthropic
- asgiref
- async-timeout
- attrs
- babel
- backoff
- beautifulsoup4
- billiard
- black
- branca
- bs4
- cachetools
- +110 more
neural-network-api-service/requirements.txt
pypi · 13- fastapi
- ipykernel
- mangum
- openai
- pandas
- pydantic
- pyngrok
- python-dotenv
- setuptools
- tiktoken
- torch
- torchvision
- uvicorn
docs/requirements.txt
pypi · 8- beautifulsoup4
- Django
- django-celery-beat
- folium
- python-dotenv
- pytz
- sphinx
- sphinx-wagtail-theme
Declared in the repository’s manifests at the indexed commit. A declared package is not proof it is used, and runtime dependencies are listed first.
This project’s features have not been analysed yet.
Export this project's context (description, README, evidence, key source files) to chat with an AI agent elsewhere.
