VerifyAI is an AI-powered tool designed to detect deepfake images, distinguishing between authentic photographs and images generated by AI models. This project utilizes a deep learning model integrated into a user-friendly web interface.
- AI vs. Real Image Classification: Classifies uploaded images as either "AUTHENTIC (REAL)" or "AI GENERATED (FAKE)".
- High Accuracy: Employs a ResNet-50 model fine-tuned on the CIFAKE dataset, achieving approximately 96% accuracy on the test set.
- Streamlit Dashboard: Provides an interactive web interface for easy image uploads and result visualization.
- (Optional) Flask API: Includes a basic Flask backend API endpoint (
/analyze) for programmatic access (currently inapi.py).
- Architecture: Transfer Learning using ResNet-50 (pre-trained on ImageNet).
- Dataset: Trained on the CIFAKE dataset ([kaggle+2] https://www.kaggle.com/datasets/birdy654/cifake-real-and-ai-generated-synthetic-images), containing 100,000 training images (50k real, 50k AI) and 20,000 test images (10k real, 10k AI).
- Performance: Achieved ~95.93% accuracy on the independent test set.
- Baseline Comparison: A simple CNN trained from scratch achieved ~86.96% accuracy but showed significant overfitting and bias.
- Input: Expects 128x128 pixel RGB images.
- Output: Binary classification (FAKE=0, REAL=1) with a confidence score.
- Explainability (Grad-CAM): Attempts to implement Grad-CAM were made but were unsuccessful due to persistent tooling/library conflicts with the nested model structure. This feature is currently omitted.
- Backend & Model: Python, TensorFlow/Keras
- Web Framework (Dashboard): Streamlit
- Web Framework (API): Flask (Optional, in
api.py) - Image Processing: Pillow, OpenCV (for dashboard visualization)
- Numerical Computing: NumPy
- Data Visualization: Matplotlib (used during development/evaluation)
-
Clone the repository:
git clone [https://github.com/Hussny-06/VerifyAI.git](https://github.com/Hussny-06/VerifyAI.git) cd VerifyAI -
Create and activate a virtual environment:
# Windows python -m venv venv .\venv\Scripts\activate # macOS/Linux python3 -m venv venv source venv/bin/activate
-
Install dependencies:
pip install -r requirements.txt
-
(Optional - Git LFS): Git LFS was implemented for model files, ensure it's installed (
git lfs install) and pull the large files (git lfs pull). .
There are two ways to run the application:
This provides the interactive web UI.
-
Ensure your virtual environment is active.
-
Run the following command from the project root directory:
streamlit run app.py
-
Open your web browser and navigate to the local URL provided (usually
http://localhost:8501). -
Use the sidebar to navigate (currently only "Single Image" is implemented).
-
Upload an image on the "Single Image" page to get a prediction.
This runs a simple API server. Note: Ensure you have renamed the original Flask app file to api.py as discussed.
-
Ensure your virtual environment is active.
-
Run the following command from the project root directory:
python api.py
-
The API will be available at
http://localhost:5000. You can send POST requests to the/analyzeendpoint with an image file attached (key: 'image').
VerifyAI/
├── .git/ # Git metadata
├── .venv/ # Virtual environment (ignored by git)
├── frontend/ # Optional Flask frontend (HTML/CSS/JS)
│ ├── index.html
│ ├── script.js
│ └── style.css
├── pages/ # Streamlit pages
│ └── 1_Single_Image.py
├── .gitignore
├── api.py # Optional Flask API
├── app.py # Streamlit entrypoint
├── baseline_cnn_v1.h5 # Saved baseline model (reference)
├── LICENSE
├── README.md
├── requirements.txt
├── resnet50_v1.keras # Trained model file
└── tl_history.npy # Training history (numpy)
- Implement Video Analysis: Extend the detection capabilities to analyze video files, potentially incorporating temporal analysis techniques (e.g., using CNN+LSTM or Transformers) to identify inconsistencies between frames.
- Enhance Streamlit Dashboard:
- Implement the remaining pages (Batch Processing, Model Performance, Dataset Explorer, About Model) as outlined in the dashboard plan.
- Add video upload and analysis functionality to the dashboard.
- Explainability: Re-attempt Grad-CAM integration using alternative methods or libraries if they become more stable, or explore other XAI techniques (like LIME or SHAP).
- Model Exploration: Experiment with other high-performing architectures like EfficientNet or Xception.
- Dataset Expansion: Augment the training data with images/videos from newer or more diverse generative models.
- Containerization: Package the application (either Flask API or Streamlit dashboard) using Docker for easier deployment and portability.
- Cloud Deployment: Deploy the application to a cloud platform (e.g., Streamlit Community Cloud, Heroku, AWS, Google Cloud).