Data & Automation Engineer • Platform Owner & Maintainer

Hi, I'm Rishav Kumar

Data & Automation Engineer | Platform Owner & Maintainer (iMigrator) β€’ Cognizant

Results-driven Data & Automation Engineer (Programmer Analyst at Cognizant) owning and maintaining enterprise data validation platforms, CI/CD automation, and test engineering utilities. Platform Owner and Maintainer of iMigrator across multiple projectsβ€”stabilising the inherited core engine and independently introducing DuckDB-based validation to achieve up to 3x throughput gains in benchmark runs and eliminate crashes on low-RAM systems. Co-developed GenAI failure diagnostics (AWS Bedrock / Claude) and Snowflake pushdown reconciliation, built standalone self-updaters and packaging, and improved Jenkins CI/CD automation across multiple project workflows.

Executive Technical Snapshot
Platform Owner
Rishav Kumar
⚑ Cognizant β€’ Automation Team

Rishav Kumar

πŸ“ Kolkata, India

Platform Owner • iMigrator
LANGUAGES & CORE TECH
Python 3.12 SQL Bash
HIGH-PERFORMANCE ENGINES
Polars LazyFrames DuckDB In-Memory Pandas PyArrow psutil (Dynamic RAM)
AI DEVELOPER TOOLS (IDE & CHAT)
Google Antigravity Claude Code (CLI) GitHub Copilot Agentic Workflows
AI MODELS & SYSTEMS
AWS Bedrock (Claude Sonnet) Google Gemini API Google ADK Prompt Engineering Agentic Workflows
DATA PLATFORMS SUPPORTED (iMIGRATOR)
Microsoft Fabric Snowflake Databricks DynamoDB Oracle PostgreSQL MS SQL Server IBM DB2 & AS400 Sybase MongoDB 15+ Heterogeneous Connectors
DEVOPS, PACKAGING & QA
Jenkins CI/CD (Groovy) PyInstaller (3-Tier Hot-Swap) Agile Lifecycle Automation (Rally) Git / GitHub
15+
Heterogeneous Database Connectors
Up to 3x
Faster Reconciliation via Polars & DuckDB
50M+
Rows Benchmark / Zero OOM Memory Allocation
12+
Data Formats & File Types Supported
// TECHNICAL ARSENAL

Core Technical Competencies

Engineered for high-throughput distributed computation, heterogeneous cloud databases, GenAI pipelines, and autonomous test harnesses.

πŸ’»

Programming & Core Tech

Production-grade systems programming, AST parsing, high-concurrency multiprocessing, and Windows OS API automation.

Python 3.12 (OOP & AST) SQL (Complex DDL / DML) Bash / Shell Multiprocessing & Concurrency Windows API Process Handling PyInstaller Packaging
⚑

High-Performance Data Engines

Memory-optimized columnar computation engines replacing traditional memory-bound iterations with zero-OOM processing for 50M+ rows.

Polars LazyFrames DuckDB In-Memory Pandas PyArrow psutil (Dynamic RAM Sizing)
πŸ› οΈ

AI Developer Tools & Workflows

Deep daily integration of frontier AI coding tools, autonomous agents, and conversational intelligence across VS Code and terminal workflows.

Google Antigravity Anthropic Claude Code (CLI) GitHub Copilot Agentic Pair Programming
☁️

Database Connectors & Cloud (15+)

Multi-platform integration adapters developed for iMigrator with native cross-database staging and temporary table joins.

Snowflake Microsoft Fabric Databricks AWS DynamoDB Oracle & Exadata PostgreSQL Amazon Aurora (PostgreSQL / MySQL) MS SQL Server (MSSQL) MySQL IBM DB2 & DB2 AS400 Sybase (jConnect & jTDS) MongoDB Amazon Redshift Amazon Athena & AWS S3 IBM Netezza SAP HANA SAS SQLite Cross-DB Staging Tables
πŸ“‚

File Formats & Extraction (12+)

High-throughput extraction adapters handling big data serializations, enterprise nested structures, and legacy enterprise formats.

Apache Parquet Apache Avro Nested JSON (Deep Flattening) XML / XSD Excel (OpenPyXL / XlsxWriter) CSV / Delimited DAT Files Flat Files Fixed-Width Formats CLOB / BLOB Extraction SAS7BDAT PDF Extraction CTRL Control Files
πŸ€–

Generative AI & Diagnostics

Autonomous failure clustering and root-cause diagnostics integrating cloud foundational models for automated reconciliation.

AWS Bedrock Anthropic Claude Sonnet Google Gemini API Google Agent Development Kit (ADK) Prompt Engineering Structured JSON Extraction
πŸš€

CI/CD & Enterprise Packaging

Collision-safe distributed test runners, atomic self-updating desktop distribution, and robust execution telemetry.

Jenkins Parameterized Pipelines Groovy Scripting Pipeline Concurrency PyInstaller (3-Tier Hot-Swap) Windows API Handling Enterprise Network Distribution
πŸ›‘οΈ

Test Governance & Data Validation

Enterprise QA orchestration, automated defect lifecycle management, DDL schema drift tracking, and Informatica pipelines.

Agile Test Management (Rally) Schema Drift Detection STTM & Data Vault Validation MD5 Keyless Reconciliation Informatica PowerCenter DML Execution Blockers
// CAREER TRAJECTORY

Professional Experience

Proven product ownership, architectural migrations, and production-grade engineering at Cognizant.

Experience at Cognizant

Cognizant Technology Solutions • Full-Time & Internship
Intern → Trainee → Programmer Analyst 03/2025 – Present
Programmer Analyst
Data & Automation Engineer • Platform Owner & Maintainer (iMigrator)
β˜… Promoted Aug 2026 08/2026 – Present
  • Platform Ownership & Releases: Took ownership of iMigrator as sole platform owner and maintainer; leading its maintenance, releases, roadmap, and user support across multiple projects, currently delivering the next major release.
  • Core Engine & DuckDB Validation: Inherited an unstable engine and independently introduced DuckDB-based validation to deliver up to 3x faster processing in benchmark runs while eliminating out-of-memory crashes on low-RAM user systems.
  • Chunked Processing & Memory Stability: Improved and stabilised inherited chunked processing logic and dynamic memory controls, resolving critical execution issues and ensuring reliable processing across high-volume datasets.
  • Snowflake Pushdown Reconciliation: Co-developed cross-database pushdown reconciliation staging heterogeneous source data into temporary Snowflake tables, executing distributed in-warehouse joins to eliminate client-side memory constraints and heavy network transfer.
  • GenAI Failure Diagnostics (AWS Bedrock & Claude): Co-developed an autonomous diagnostic module clustering validation failure patterns and generating structured JSON root-cause classifications and remediation SQL.
  • Connectors & User Support: Built and extended connector support for MongoDB, Snowflake (token-based authentication), CTRL files, and DB2; actively support users across all 15+ connectors, updating connection and validation code as issues arise.
  • Self-Updater & Packaging: Built standalone self-updater utilities and production executable packaging using PyInstaller, enabling automated distribution and updates with rollback safeguards.
  • Jenkins CI/CD Automation: Improved an early test-phase pipeline, brought it live into active use across multiple projects, and extended Jenkins automation to additional project scripts beyond iMigrator.
  • Rally Test Automation & Result Sharing: Built standalone automation utilities for bulk test case creation and execution updates in Rally; created a two-part executable design allowing scripts to write JSON results that are loaded into PostgreSQL without distributing database credentials.
Programmer Analyst Trainee
ETL Testing & Data Quality Specialist → Automation Developer
β˜… Top Performance Rating 08/2025 – 08/2026
Phase 1: ETL Testing (First 7 Months) Phase 2: Automation Developer (Automation Team)
  • Top Performance Rating & Accelerated Promotion: Conferred top performance rating in 1st-year confirmation appraisal; recognized by leadership for rapid engineering mastery and pivotal architecture contributions to the iMigrator platform, earning accelerated promotion to Programmer Analyst in August 2026.
  • Automation Team Transition (Subsequent Months): Following the initial 7-month ETL testing phase, transitioned into the Automation Team as an Automation Developer; authored modular Python utilities for automated test case generation, data-stat profiling, defect verification, and foundational reconciliation routines for iMigrator.
  • Enterprise ETL Testing (First 7 Months): Executed comprehensive end-to-end source-to-target test verification, data quality audits, and migration testing across multi-LOB data warehouse pipelines.
  • SQL & Reconciliation Verification: Formulated complex SQL reconciliation queries across Snowflake and Oracle targets, validating business transformations, primary key constraints, null bounds, and precision.
  • Defect Lifecycle Management: Documented and tracked critical ETL pipeline anomalies, schema mismatches, and data drift in CA Agile Central (Rally), collaborating closely with data engineering teams to remediate pipeline bugs prior to production cuts.
Data Engineering Intern
Informatica PowerCenter & ETL Foundations
03/2025 – 07/2025
  • Healthcare Data Integration Pipeline: Designed and implemented an end-to-end ETL integration system using Informatica PowerCenter to process multi-source healthcare insurance feeds (Group, Subgroup, Subscriber fixed-width and comma-delimited flat files) into dimensional Oracle tables.
  • Transformations & Data Cleansing: Built robust mapping logic leveraging Expression, Filter, Router, Joiner, and Sequence Generator transformations to standardize formats, eliminate duplicate records, and route error data to dedicated exception tables.
  • XML Generation & XSD Validation: Engineered a downstream publishing workflow converting transformed subscriber data into structured, personalized XML welcome letters complying with strict XSD schema validation standards.
  • Workflow Design & Scheduling: Created, configured, and monitored automated workflow sessions in Informatica Workflow Manager, analyzing session logs and performance metrics to optimize throughput.
// PRODUCTION IMPACT

Featured Engineering Projects

Core platforms, autonomous assistants, and enterprise distribution systems engineered for high scale.

iMigrator β€” Data Validation & Reconciliation Platform

Python 3.12 β€’ DuckDB β€’ Snowflake Pushdown β€’ Multi-DB Connectors (15+ Platforms) β€’ AWS Bedrock β€’ Claude β€’ Jenkins

High-throughput enterprise data reconciliation platform adopted across multiple projects. Stabilised the inherited core pipeline and independently introduced DuckDB-based validation, delivering up to 3x processing speed in benchmark runs and eliminating crashes on low-RAM systems.

β€’ Platform Ownership & Connectors: Own and maintain the platform across multiple projects; built and extended support for MongoDB, Snowflake (token-based auth), CTRL files, and DB2 while supporting users across 15+ connectors.

β€’ In-Warehouse Pushdown Reconciliation: Co-developed cross-database pushdown reconciliation staging heterogeneous source data into temporary Snowflake tables for distributed in-warehouse joins.

β€’ Chunked Processing & Stability: Improved and stabilised inherited chunked processing routines and dynamic memory controls, eliminating crashes on user workstations.

β€’ Keyless Fallback & Safety: Owned and maintained keyless fallback reconciliation routines and automated DML execution safeguards.

β€’ GenAI Failure Diagnostics: Co-developed autonomous diagnostics with AWS Bedrock / Claude Sonnet, clustering mismatch patterns and returning structured JSON root-cause classifications.

β€’ Reporting & Packaging: Improved HTML validation reports with enhanced console logging and fallbacks; built standalone self-updaters and packaging via PyInstaller.

DataCraft β€” Autonomous SQL & Data Engineering Assistant

Python β€’ Google Agent Development Kit (ADK) β€’ Gemini API β€’ Firestore β€’ Snowflake β€’ BigQuery β€’ PostgreSQL β€’ Vertex AI

Autonomous multi-modal data engineering agent built with the Google ADK during Google's 4-Hour "Build with Gemini" hackathon, earning the official Google Developer Badge & Credly Certification.

β€’ Dynamic Schema Catalog: Automated Firestore schema registration cataloging database schemas, primary keys, and data types across heterogeneous sources (list_tables, get_table_details, add_table).

β€’ Anti-Pattern Detection: AST parsing with sqlparse to detect full table scans, missing filters, and uncapped sorting with Snowflake, BigQuery, and PostgreSQL optimizations.

β€’ Multi-Modal Architecture Generation: Produced visual ER diagrams and animated Kafka event-streaming architecture diagrams leveraging Gemini multimodal models on Vertex AI with GCS storage.

β€’ Vertex AI Memory Bank: Integrated PreloadMemoryTool for session-level dialect persistence, paired with AgentEngineSandboxCodeExecutor and a responsive FastAPI/A2UI card interface.

Enterprise Packaging & Multi-Tier Self-Updater

Python β€’ PyInstaller β€’ Windows OS Architecture β€’ SQLite β€’ Enterprise Network Distribution

Architected a zero-downtime, atomic client updater (Launcher → Updater → Production Runtime) distributing updates across enterprise client installations with automated build verification and zero-downtime hot-swap.

β€’ Atomic Hot-Swap: Decoupled process execution so the updater replaces running binaries without file-lock collisions, backed by rollback protection.

β€’ Zero Console Flashing: Leveraged Windows API process handling to provide silent execution with clean SQLite telemetry logging and network distribution.

β€’ CLI Maintenance Utility: Authored automated CLI diagnostics and repair utilities for rapid environment health checks and client distribution audits.

Healthcare Payer Data Integration Pipeline

Informatica PowerCenter β€’ Oracle SQL β€’ XML / XSD β€’ Flat File Processing

Production ETL integration pipeline processing fixed-width and comma-delimited healthcare feeds into dimensional Oracle tables with automated XML welcome letter distribution.

β€’ Multi-Feed Processing: Cleaned, validated, and normalized multi-tier patient datasets across complex business transformations.

β€’ Schema Compliance: Validated outgoing XML welcome records against strict enterprise XSD schema definitions.

// VERIFIED RECOGNITION

Credentials & Education

Verified industry credentials, certifications, and academic foundations.

πŸ†

Build with Gemini (Track 3) β€” Software Developer

Google Developers β€’ Issued Feb 2026
View Google Badge β†—  β€’  Verify on Credly β†—
πŸŽ–οΈ

Context Engineering Foundation

Cognizant β€’ RAG & AI Systems β€’ Issued Jan 2026
Verify on Credly β†—
πŸŽ“

B.Tech in Computer Science & Engineering (AI)

Noida Institute of Engineering & Technology (NIET), Greater Noida β€’ 2021 – 2025

Ready to Accelerate Your Data Architecture?

Specializing in Data Engineering, High-Throughput Reconciliation, ETL Automation, and AI-Driven Data Systems. Let's discuss modern data pipelines, Polars, DuckDB, or generative AI architecture.