Gen AI Program Course

90 Hours

CORE

24/7

AI Course Group

  1. Introduction to AI and ML Concepts (4 hours)
  • Overview of AI
    • Definition, history of AI
    • Data Science, AI, and ML
    • AI in everyday Industry applications (virtual assistants, chatbots, etc.)
  • Roles and responsibilities in Data Science and AI
  • Types of AI Tasks
    • Regression vs Classification vs Clustering
    • Supervised, Unsupervised, Semi-supervised, and Reinforcement learning
  • Core Machine Learning Concepts
    • Training, testing, validation
  1. Statistics and Probability for Machine Learning (4 hours)
  • Descriptive Statistics
    • Measures of central tendency (mean, median, mode)
    • Measures of dispersion (variance, standard deviation)
  • Probability Basics
    • Basic probability rules, conditional probability
    • Bayes’ theorem and its applications in ML
  • Hypothesis Testing
    • Null vs alternative hypothesis
    • p-value, significance level, and confidence intervals
  • Correlation
    • Types of Correlation
    • Correlation Coefficient
    • Visualization of Correlation
  1. Advanced Exploratory Data Analysis (EDA) (4 hours) 
  • Data Cleaning
    • Identifying and handling missing data (mean imputation, drop missing)
    • Removing outliers (IQR, Z-score method)
  • Data Preprocessing
    • Normalization and Standardization
    • Encoding categorical variables
  • Visualization Techniques
    • Bar plots, histograms, box plots, pair plots
  1. Introduction to Machine Learning Algorithms (8 hours)
  • Regression Algorithms
    • Single Linear Regression, Multi Linear Regression
  • Classification Algorithms
    • Logistic Regression: Decision boundary, sigmoid function
  • Clustering Algorithms
    • K-Means Clustering and Elbow method for optimal clusters
  • Other Key Algorithms
    • KNN (K Nearest Neighbours) , SVM (Support Vector Machine)
    • Ensemble methods
  1. Introduction to Neural Networks (4 hours)
  • Artificial Neurons and Layers
    • Structure of a neuron, activation functions (ReLU, Sigmoid, Tanh)
  • Feedforward Networks
    • Single-layer vs multi-layer networks
    • Forward propagation, Backpropagation
  • Simple Classification with Neural Networks
    • Building a simple neural network using TensorFlow/Keras
  1. Data Visualization and Communication (2 hrs)
    • Importance of data visualization
    • Tools: Matplotlib, Seaborn, Tableau
    • Communicating results and insights effectively
    • Hands-on: Creating visualizations from datasets
  1. Data Engineering Fundamentals (2 hours)
  • Data Pipelines
    • Understanding ETL (Extract, Transform, Load)
    • Data processing: Batch vs real-time processing
  • Basic SQL
    • SELECT, JOIN, GROUP BY, and aggregate functions (COUNT, SUM, etc.)
    • Query optimization and indexing
  1. Capstone Project (3 hrs)
    • Project: Integrate skills to solve a real-world problem
    • Presenting findings and model performance
    • Peer reviews and feedback sessions
  1. Introduction to AI and ML Concepts (4 hours)
  • Overview of AI
    • Definition, history of AI
    • Data Science, AI, and ML
    • AI in everyday Industry applications (virtual assistants, chatbots, etc.)
  • Roles and responsibilities in Data Science and AI
  • Types of AI Tasks
    • Regression vs Classification vs Clustering
    • Supervised, Unsupervised, Semi-supervised, and Reinforcement learning
  • Core Machine Learning Concepts
    • Training, testing, validation
  1. Statistics and Probability for Machine Learning (6 hours)
  • Descriptive Statistics
    • Measures of central tendency (mean, median, mode)
    • Measures of dispersion (variance, standard deviation)
  • Probability Basics
    • Basic probability rules, conditional probability
    • Bayes’ theorem and its applications in ML
  • Hypothesis Testing
    • Null vs alternative hypothesis
    • p-value, significance level, and confidence intervals
  • Correlation
    • Types of Correlation
    • Correlation Coefficient
    • Visualization of Correlation
  • Data Distributions
    • Understanding frequency distributions
    • Skewness and Kurtosis: Shape of distributions
    • Histograms and probability density functions
  • Statistical Tests
    • t-Test: Independent and paired samples
    • ANOVA (Analysis of Variance): Comparing means across groups
    • Chi-square test: Categorical data relationships
    • P-value interpretation and statistical significance
  1. Exploratory Data Analysis (EDA) (8 hours)

3.1 Data processing (6 hours)

  • Data Cleaning
    • Identifying and handling missing data (mean imputation, drop missing)
    • Removing outliers (IQR, Z-score method)
  • Data Preprocessing
    • Normalization and Standardization
    • Encoding categorical variables
  • Visualization Techniques
    • Bar plots, histograms, box plots, pair plots
  • Data Imbalance
    • Oversampling (SMOTE) and under sampling techniques for balanced datasets

3.2 Univariate and Bivariate Analysis (2 hours)

  • Univariate Analysis
    • Distribution of a single variable
    • Visualization: Histograms, KDE plots, Box plots
  • Bivariate Analysis
    • Relationships between two variables (numeric-numeric, numeric-categorical, categorical-categorical)
    • Scatter plots, bar plots, pair plots
    • Correlation analysis for numeric-numeric data
    • Cross-tabulation and Chi-square tests for categorical-categorical relationships
  1. Machine Learning Algorithms (10 hours)

4.1 ML Algorithm and Model Building (6 hours)

  • Regression Algorithms
    • Single Linear Regression, Multi Linear Regression
  • Classification Algorithms
    • Logistic Regression: Decision boundary, sigmoid function
  • Clustering Algorithms
    • K-Means Clustering and Elbow method for optimal clusters
  • Other Advanced Algorithms
    • KNN (K Nearest Neighbours) , SVM (Support Vector Machine)
    • Ensemble methods:
      • Bagging: Random Forest algorithm
      • Boosting: Gradient Boosting, AdaBoost, XGBoost

4.2 ML Fine tuning and Testing (4 hours)

  • Model Hyperparameter tuning
    • Cross-validation, K-fold, and Leave-one-out Cross-validation
    • Hyperparameter tuning (GridSearchCV, RandomizedSearchCV)
  • Model Evaluation
    • Regression:
      • MSE
      • MAE
      • RMSE
      • R2 and Adjusted R2
    • Classification:
      • Accuracy,
      • Precision,
      • Recall,
      • F1 Score
  1. Data Science Project Life Cycle (2 hours)
  • Problem Definition
    • Defining business problems and translating into ML tasks
    • Understanding stakeholders’ requirements
  • Data Acquisition and Feature Engineering
    • Gathering data from APIs, databases, and scraping
    • Feature selection and dimensionality reduction (PCA)
  • Model Selection
    • Choosing between algorithms based on problem type (classification, regression)
    • Ensemble models vs standalone models
  1. Deep Learning and Neural Networks (12 hours)

6.1 Fundamentals (2 hours)

    • Overview of Neural Network
    • Forward propagation, Backpropagation
    • Input layers, Hidden Layers, Output Layers
    • Optimization, Activation Functions

6.2 Types of Deep Learning Usecases (6 hours)

  • Convolutional Neural Networks (CNNs)
    • Convolution operations, filters, feature maps
    • Pooling layers, fully connected layers
    • Applications in image classification (MNIST, CIFAR-10)
  • Recurrent Neural Networks (RNNs)
    • Sequence data handling, vanishing gradients problem
    • Long Short-Term Memory (LSTM), Gated Recurrent Units (GRUs)
    • Applications in time-series prediction, NLP
  • LSTM Architecture
    • Gates in LSTM: Forget gate, input gate, and output gate
    • Usecase and practical application

6.3 Computer Vision (4 hours)

  • Image Processing
    • Image transformations (rotation, resizing, cropping)
    • Edge detection (Canny, Sobel filters)
  • Object Detection
    • YOLO (You Only Look Once), SSD (Single Shot Multibox Detector)
    • Applications in autonomous driving, security, and robotics
  • Image Classification
    • Pretrained models (ResNet, VGG) and transfer learning
  1. Natural Language Processing (NLP) (5 hours)
  • Text Preprocessing
    • Tokenization, Lemmatization, Stemming
    • Stopwords removal, handling punctuation and special characters
  • Word Embeddings
    • Word2Vec, Skip-gram and Continuous Bag of Words (CBOW)
    • GloVe and its advantages
  • Text Classification and Sentiment Analysis
    • TF-IDF, Bag of Words, and embedding-based approaches
    • Sentiment Analysis with pre-trained models
  • Named Entity Recognition (NER)
    • Techniques for extracting entities (names, dates, locations) from text
    • Pretrained models (Spacy, HuggingFace)
  1. Data Engineering and Big Data Tools (3 hours)
  • Data Warehousing
    • ETL processes and tools (Apache Nifi, Talend)
    • Data Lakes vs Data Warehouses (Hadoop, AWS S3, Redshift)
  • Big Data Processing
    • Introduction to Hadoop Ecosystem (HDFS, MapReduce, Hive)
    • Real-time processing with Apache Spark (RDD, DataFrames, SparkSQL)
  • NoSQL Databases
    • MongoDB, Cassandra: Basics, use cases, and scaling strategies
  1. MLOps (3 hours)
  • Model Deployment
    • Docker containers for ML model encapsulation
    • Kubernetes for container orchestration
    • Deployment options: Cloud-based (AWS SageMaker, GCP AI Platform) vs On-prem
  • Continuous Integration/Continuous Deployment (CI/CD)
    • GitLab CI, Jenkins, and automated pipeline setup
    • Automated testing for models
  • Model Monitoring and Retraining
    • Performance drift, data drift monitoring tools
    • Automating model retraining with Airflow
  1. Data Visualization (3 hours)
  • Tools for Visualization
    • Power BI, Tableau for interactive dashboards
    • Integrating ML models into dashboards
  • Advanced Techniques
    • Heatmaps, correlation matrices, geospatial visualizations
    • Visualization best practices for large datasets
  1. Capstone Project (4 hrs)
    • Project: Integrate skills to solve a real-world problem
    • Presenting findings and model performance
    • Peer reviews and feedback sessions
  1. Introduction to AI and ML Concepts (4 hours)
  • Overview of AI
    • Definition, history of AI
    • Data Science, AI, and ML
    • AI in everyday Industry applications (virtual assistants, chatbots, etc.)
  • Roles and responsibilities in Data Science and AI
  • Types of AI Tasks
    • Regression vs Classification vs Clustering
    • Supervised, Unsupervised, Semi-supervised, and Reinforcement learning
  • Core Machine Learning Concepts
    • Training, testing, validation
  1. Statistics and Probability for Machine Learning (8 hours)
  • Descriptive Statistics
    • Measures of central tendency (mean, median, mode)
    • Measures of dispersion (variance, standard deviation)
  • Probability Basics
    • Basic probability rules, conditional probability
    • Bayes’ theorem and its applications in ML
  • Bayesian Statistics
    • Bayes’ Theorem and its applications in probabilistic models
    • Prior, posterior, and likelihood in Bayesian inference
    • Applications in decision-making (spam filtering, A/B testing)
  • Hypothesis Testing
    • Null vs alternative hypothesis
    • p-value, significance level, and confidence intervals
  • Correlation
    • Types of Correlation
    • Correlation Coefficient
    • Visualization of Correlation
  • Data Distributions
    • Understanding frequency distributions
    • Skewness and Kurtosis: Shape of distributions
    • Histograms and probability density functions
  • Statistical Tests
    • t-Test: Independent and paired samples
    • ANOVA (Analysis of Variance): Comparing means across groups
    • Chi-square test: Categorical data relationships
    • P-value interpretation and statistical significance
  1. Exploratory Data Analysis (EDA) (10 hours)

3.1 Data processing (6 hours)

  • Data Cleaning
    • Identifying and handling missing data (mean imputation, drop missing)
    • Removing outliers (IQR, Z-score method)
  • Data Preprocessing
    • Normalization and Standardization
    • Encoding categorical variables
  • Visualization Techniques
    • Bar plots, histograms, box plots, pair plots
  • Data Imbalance
    • Oversampling (SMOTE) and under sampling techniques for balanced datasets

3.2 Univariate and Bivariate Analysis (2 hours)

  • Univariate Analysis
    • Distribution of a single variable
    • Visualization: Histograms, KDE plots, Box plots
  • Bivariate Analysis
    • Relationships between two variables (numeric-numeric, numeric-categorical, categorical-categorical)
    • Scatter plots, bar plots, pair plots
    • Correlation analysis for numeric-numeric data
    • Cross-tabulation and Chi-square tests for categorical-categorical relationships

3.3 Data Visualization and Storytelling (2 hours)

  • Advanced Visualization Tools
    • Interactive dashboards with Plotly, Bokeh
    • Visualizing complex relationships with Seaborn pair plots
  •  Effective Communication of Data Insights
    • Best practices for presenting data to non-technical audiences
    • Data storytelling techniques: Choosing the right chart types
    • Integrating data visualizations into reports and dashboards (Power BI, Tableau)

 

  1. Machine Learning Algorithms (14 hours)

4.1 ML Algorithm and Model Building (8 hours)

  • Regression Algorithms
    • Single Linear Regression, Multi Linear Regression
  • Classification Algorithms
    • Logistic Regression: Decision boundary, sigmoid function
  • Clustering Algorithms
    • K-Means Clustering and Elbow method for optimal clusters
  • Other Advanced Algorithms
    • KNN (K Nearest Neighbours) , SVM (Support Vector Machine)
    • Ensemble methods:
      • Bagging: Random Forest algorithm
      • Boosting: Gradient Boosting, AdaBoost, XGBoost

4.2 ML Fine tuning and Testing (4 hours)

  • Model Hyperparameter tuning
    • Cross-validation, K-fold, and Leave-one-out Cross-validation
    • Hyperparameter tuning (GridSearchCV, RandomizedSearchCV)
  • Model Evaluation
    • Regression:
      • MSE
      • MAE
      • RMSE
      • R2 and Adjusted R2
    • Classification:
      • Accuracy,
      • Precision,
      • Recall,
      • F1 Score

4.3 Advanced Regression Techniques (2 hours)

  • Regularization Techniques
    • Ridge and Lasso regression for handling multicollinearity and overfitting
    • ElasticNet: Combining Lasso and Ridge for balanced regularization
  1. Data Science Project Life Cycle (4 hours)
  • Problem Definition
    • Defining business problems and translating into ML tasks
    • Understanding stakeholders’ requirements
  • Data Acquisition and Feature Engineering
    • Gathering data from APIs, databases, and scraping
    • Feature selection and dimensionality reduction (PCA)
  • Model Selection
    • Choosing between algorithms based on problem type (classification, regression)
    • Ensemble models vs standalone models
  1. Deep Learning and Neural Networks (16 hours)

6.1 Fundamentals (2 hours)

    • Overview of Neural Network
    • Forward propagation, Backpropagation
    • Input layers, Hidden Layers, Output Layers
    • Optimization, Activation Functions

6.2 Types of Deep Learning Usecases (8 hours)

  • Convolutional Neural Networks (CNNs)
    • Convolution operations, filters, feature maps
    • Pooling layers, fully connected layers
    • Applications in image classification (MNIST, CIFAR-10)
  • Recurrent Neural Networks (RNNs)
    • Sequence data handling, vanishing gradients problem
    • Long Short-Term Memory (LSTM), Gated Recurrent Units (GRUs)
    • Applications in time-series prediction, NLP
  • LSTM Architecture
    • Gates in LSTM: Forget gate, input gate, and output gate
    • Usecase and practical application
  • Transformer Models
    • Transformer architecture for sequence processing
    • BERT (Bidirectional Encoder Representations from Transformers)
    • Applications in NLP: Machine translation, text classification

6.3 Advanced Computer Vision Techniques (6 hours)

  • Image Processing
    • Image transformations (rotation, resizing, cropping)
    • Edge detection (Canny, Sobel filters)
    • Techniques for pixel-level image understanding
  • Image Classification
    • Pretrained models (ResNet, VGG) and transfer learning
  • 3D Computer Vision
    • Depth estimation, point cloud analysis
  • Object Detection
    • YOLO (You Only Look Once), SSD (Single Shot Multibox Detector)
    • Applications in autonomous driving, security, and robotics
  • Real-Time Object Detection
    • Real-time applications with YOLOv8 and Faster R-CNN
  1. Introduction to Time Series Forecasting (4 hours)
  • Time Series Data Overview
  • Statistical Concepts in Forecasting
  • Moving Average (MA) and Exponential Smoothing Methods
  • Stationarity and Differencing
  • ARIMA and SARIMA Models and Advance ARIMA family models
  • Random Forest and Gradient Boosting for Time Series (6 hours)
  1. Natural Language Processing (NLP) (6 hours)
  • Text Preprocessing
    • Tokenization, Lemmatization, Stemming
    • Stopwords removal, handling punctuation and special characters
  • Word Embeddings
    • Word2Vec, Skip-gram and Continuous Bag of Words (CBOW)
    • GloVe and its advantages
  • Text Classification and Sentiment Analysis
    • TF-IDF, Bag of Words, and embedding-based approaches
    • Sentiment Analysis with pre-trained models
  • Named Entity Recognition (NER)
    • Techniques for extracting entities (names, dates, locations) from text
    • Pretrained models (Spacy, HuggingFace)
  1. Data Engineering and Big Data Tools (4 hours)
  • Data Warehousing
    • ETL processes and tools (Apache Nifi, Talend)
    • Data Lakes vs Data Warehouses (Hadoop, AWS S3, Redshift)
  • Big Data Processing
    • Introduction to Hadoop Ecosystem (HDFS, MapReduce, Hive)
    • Real-time processing with Apache Spark (RDD, DataFrames, SparkSQL)
  • NoSQL Databases
    • MongoDB, Cassandra: Basics, use cases, and scaling strategies
  1. MLOps (5 hours)
  • Model Deployment
    • Docker containers for ML model encapsulation
    • Kubernetes for container orchestration
    • Deployment options: Cloud-based (AWS SageMaker, GCP AI Platform) vs On-prem
  • Model Monitoring and Retraining
    • Performance drift, data drift monitoring tools
    • Automating model retraining with Airflow
  • MLOps in MLFLow
    • MLflow Tracking
    • MLflow Projects
    • MLflow Models
    • MLflow Model Registry
    • Advanced MLOps Techniques with MLflow
  1. AI for Business and Operations (3 hours)
  • Building AI Products
    • Strategies for integrating AI into business operations
    • Estimating AI project ROI
  • AI Operations
    • Managing AI systems in production, scaling AI solutions
    • Cloud vs On-prem AI infrastructure decisions
  1. Specialized Tracks – Training (6 hours)
  • Track 1: Advanced NLP Engineer
    • Building chatbots, machine translation systems
    • Large-scale language model training
  • Track 2: Advanced Computer Vision Engineer
    • Real-time CV solutions for drones, surveillance, and robotics
    • 3D vision and augmented reality (AR) applications
  • Track 3: Advanced Data Scientist
    • Advanced Bayesian methods, probabilistic programming
    • Modeling uncertainty and risk in data science
  • Track 4: Advanced MLOps Specialist
    • Fully automated pipeline deployments, auto-scaling
    • Monitoring and managing thousands of models at scale
  1. Capstone Project (6 hrs)
    • Project: Integrate skills to solve a real-world problem
    • Presenting findings and model performance
    • Peer reviews and feedback sessions

error: Content is protected !!