Available for Data Engineer & Automation Roles

Hi, I'm Rishav Kumar

Software Development & Automation Engineer β€’ Programmer Analyst at Cognizant

Results-driven Software Development & Automation Engineer (Programmer Analyst at Cognizant) and technical product owner specializing in high-throughput data validation engines, autonomous CI/CD pipelines, and enterprise automation utilities. Primary Product Owner and Lead Developer of iMigrator scaled across 60+ enterprise projectsβ€”re-engineering core reconciliation pipelines with Polars LazyFrames and DuckDB for 3x throughput and zero-OOM execution on 50M+ row tables. Proven leadership integrating GenAI (AWS Bedrock / Anthropic Claude) for automated failure diagnostics, PyInstaller multi-tier self-updaters, and collision-safe Jenkins automation.

Executive Technical Snapshot
Lead Automation Engineer
Rishav Kumar
⚑ Cognizant FTE β€’ Automation Team

Rishav Kumar

πŸ“ Kolkata, West Bengal, India

Product Owner β€’ iMigrator Core Platform
LANGUAGES & CORE TECH
Python 3.12 SQL Bash
HIGH-PERFORMANCE ENGINES
Polars LazyFrames DuckDB In-Memory Pandas PyArrow psutil (Dynamic RAM)
AI DEVELOPER TOOLS (IDE & CHAT)
Google Antigravity Claude (Code CLI & Chat) GitHub Copilot (VS Code & Chat) OpenAI Codex
AI MODELS & SYSTEMS
AWS Bedrock (Claude Sonnet) Google Gemini API Google ADK Prompt Engineering Agentic Workflows
DATA PLATFORMS SUPPORTED (iMIGRATOR)
Microsoft Fabric Snowflake Databricks DynamoDB Oracle PostgreSQL MS SQL Server IBM DB2 & AS400 Sybase MongoDB +17 Heterogeneous Connectors
DEVOPS, PACKAGING & QA
Jenkins CI/CD (Groovy) PyInstaller (3-Tier Hot-Swap) CA Agile Central (Rally / Pyral) Git / GitHub
60+
Enterprise Projects Powered by iMigrator
3x
Faster Reconciliation via Polars & DuckDB
50M+
Rows / Zero OOM Memory Allocation
17+
Multi-Database Heterogeneous Connectors
// TECHNICAL ARSENAL

Core Technical Competencies

Engineered for high-throughput distributed computation, heterogeneous cloud databases, GenAI pipelines, and autonomous test harnesses.

πŸ’»

Programming & Core Tech

Production-grade systems programming, AST parsing, high-concurrency multiprocessing, and Windows OS API automation.

Python 3.12 (OOP & AST) SQL (Complex DDL / DML) Bash / Shell Multiprocessing & Concurrency Windows API Process Handling PyInstaller Packaging
⚑

High-Performance Data Engines

Memory-optimized columnar computation engines replacing traditional memory-bound iterations with zero-OOM processing for 50M+ rows.

Polars LazyFrames DuckDB In-Memory Pandas PyArrow Dask psutil (Dynamic RAM Sizing)
πŸ› οΈ

AI Developer Tools & Workflows

Deep daily integration of frontier AI coding tools, autonomous agents, and conversational intelligence across VS Code and terminal workflows.

Google Antigravity Anthropic Claude Code (CLI) Claude Web / Chat GitHub Copilot (VS Code) Copilot Chat OpenAI Codex Agentic Pair Programming
☁️

Database Connectors & Cloud (17+)

Multi-platform integration adapters developed for iMigrator with native cross-database staging and temporary table joins.

Snowflake Microsoft Fabric Databricks AWS DynamoDB Oracle & Exadata PostgreSQL Amazon Aurora (PostgreSQL / MySQL) MS SQL Server (MSSQL) MySQL IBM DB2 & DB2 AS400 Sybase (jConnect & jTDS) MongoDB Amazon Redshift Amazon Athena & AWS S3 IBM Netezza SAP HANA SAS SQLite Cross-DB Staging Tables
πŸ“‚

File Formats & Extraction (12+)

High-throughput extraction adapters handling big data serializations, enterprise nested structures, and legacy enterprise formats.

Apache Parquet Apache Avro Nested JSON (Deep Flattening) XML / XSD Excel (OpenPyXL / XlsxWriter) CSV / Delimited DAT Files Flat Files Fixed-Width Formats CLOB / BLOB Extraction SAS7BDAT PDF Extraction CTRL Control Files
πŸ€–

Generative AI & Diagnostics

Autonomous failure clustering and root-cause diagnostics integrating cloud foundational models for automated reconciliation.

AWS Bedrock Anthropic Claude Sonnet Google Gemini API Google Agent Development Kit (ADK) Prompt Engineering Structured JSON Extraction
πŸš€

CI/CD & Enterprise Packaging

Collision-safe distributed test runners, atomic self-updating desktop distribution, and robust execution telemetry.

Jenkins Parameterized Pipelines Groovy Scripting Lockable Resources PyInstaller (3-Tier Hot-Swap) Windows API Handling UNC Share Distribution
πŸ›‘οΈ

Test Governance & Data Validation

Enterprise QA orchestration, automated defect lifecycle management, DDL schema drift tracking, and Informatica pipelines.

CA Agile Central (Rally / Pyral) Schema Drift Detection STTM & Data Vault Validation MD5 Keyless Reconciliation Informatica PowerCenter DML Execution Blockers
// CAREER TRAJECTORY

Professional Experience

Proven product ownership, architectural migrations, and production-grade engineering at Cognizant.

Experience at Cognizant

Cognizant Technology Solutions • Full-Time & Internship
Intern → Trainee → Programmer Analyst 03/2025 – Present
Programmer Analyst
Automation Developer • Lead Automation Engineer & Product Owner (Automation Team)
β˜… Promoted 7 Aug 2026 08/2026 – Present
  • Platform Ownership & Scale: Primary Product Owner and Lead Developer of iMigrator (v6.4 → v6.6/v7), scaling validation and automated reconciliation across 60+ enterprise client projects using generic, reusable architecture.
  • High-Performance Engine Migration: Re-architected core comparison from Pandas memory-bound execution to Polars LazyFrames and DuckDB in-memory relational engine, delivering 3x faster processing and a 60%+ memory reduction.
  • Adaptive Memory & Chunk Streaming: Implemented psutil-based dynamic memory allocation (sizing to 70–85% of available RAM) and disk-backed chunk streaming (fetch_to_disk), enabling low-memory systems (8GB/16GB RAM) to process 50M+ records across 50 columns in 10–15 minutes with zero out-of-memory errors.
  • Cross-Database Join & Snowflake Pushdown Engine: Engineered a specialized cross-database reconciliation engine that stages heterogeneous source data into temporary Snowflake tables, executing distributed in-warehouse joins directly on Snowflake compute to eliminate multi-terabyte network data transfer and client-side memory limits.
  • GenAI Failure Diagnostics (AWS Bedrock & Claude): Architected an autonomous diagnostic module clustering failure patterns, sampling 3 unique records per mismatch pattern, and generating structured JSON containing root-cause classifications, remediated SQL queries, and confidence scores; decoupled architecture ensures zero impact on validation runs if AI is offline.
  • Heterogeneous Data Connectors (17+ Platforms & 12+ Formats): Engineered high-throughput extraction connectors, query adapters, and cross-database staging modules supporting heterogeneous sources/targetsβ€”including Microsoft Fabric, Snowflake, DynamoDB, Oracle, PostgreSQL, IBM DB2 & DB2 AS400, AWS Athena/Aurora, Amazon Redshift, MongoDB, MS SQL Server, Sybase, and formats like Parquet, Avro, nested JSON, XML, and legacy CTRL files.
  • Enterprise Packaging & Multi-Tier Self-Updater: Designed a 3-tier zero-downtime hot-swap updater (Launcher → Updater → Real App) using PyInstaller, Windows API process handling, and UNC share distribution with rollback protection and zero console flashing.
  • CI/CD Pipeline Orchestration: Deployed multi-script automation pipelines on Jenkins using parameterized Groovy scripts, dynamic runtime environment fetching, and Lockable Resources to prevent collisions across concurrent pipeline runs on shared network storage.
  • Rally Test Automation (Pyral): Built standalone spreadsheet-driven automation for bulk test case/step creation, result uploads, and log deletion, alongside hybrid live streaming of execution verdicts directly into Rally.
Programmer Analyst Trainee
ETL Testing & Data Quality Specialist → Automation Developer
β˜… Top 5-Star Appraisal Rating 08/2025 – 08/2026
Phase 1: ETL Testing (First 7 Months) Phase 2: Automation Developer (Automation Team)
  • Top 5-Star Appraisal Rating & Promotion: Conferred top 5-Star (“5 - High”) performance rating across all 13 core evaluation competencies in 1st-Year Confirmation Appraisal; commended by engineering leadership for rapid Python mastery, dedicated support, and pivotal contributions to the iMigrator module, accelerating promotion to Programmer Analyst on August 7, 2026.
  • Automation Team Transition (Subsequent Months): Following the initial 7-month ETL testing phase, transitioned into the Automation Team as an Automation Developer; authored modular Python utilities for automated test case generation, data-stat profiling, defect verification, and foundational reconciliation routines for iMigrator.
  • Enterprise ETL Testing (First 7 Months): Executed comprehensive end-to-end source-to-target test verification, data quality audits, and migration testing across multi-LOB data warehouse pipelines.
  • SQL & Reconciliation Verification: Formulated complex SQL reconciliation queries across Snowflake and Oracle targets, validating business transformations, primary key constraints, null bounds, and precision.
  • Defect Lifecycle Management: Documented and tracked critical ETL pipeline anomalies, schema mismatches, and data drift in CA Agile Central (Rally), collaborating closely with data engineering teams to remediate pipeline bugs prior to production cuts.
Data Engineering Intern
Informatica PowerCenter & ETL Foundations
03/2025 – 07/2025
  • Healthcare Data Integration Pipeline: Designed and implemented an end-to-end ETL integration system using Informatica PowerCenter to process multi-source healthcare insurance feeds (Group, Subgroup, Subscriber fixed-width and comma-delimited flat files) into dimensional Oracle tables.
  • Transformations & Data Cleansing: Built robust mapping logic leveraging Expression, Filter, Router, Joiner, and Sequence Generator transformations to standardize formats, eliminate duplicate records, and route error data to dedicated exception tables.
  • XML Generation & XSD Validation: Engineered a downstream publishing workflow converting transformed subscriber data into structured, personalized XML welcome letters complying with strict XSD schema validation standards.
  • Workflow Design & Scheduling: Created, configured, and monitored automated workflow sessions in Informatica Workflow Manager, analyzing session logs and performance metrics to optimize throughput.
// PRODUCTION IMPACT

Featured Engineering Projects

Core platforms, autonomous assistants, and enterprise distribution systems engineered for high scale.

iMigrator Platform Evolution (v6.4 → v6.6/v7)

Python 3.12 β€’ Polars β€’ DuckDB β€’ Snowflake Pushdown β€’ Multi-DB Connectors (17+ Platforms) β€’ AWS Bedrock β€’ Claude β€’ Jenkins

High-throughput enterprise data reconciliation platform scaled across 60+ projects. Re-engineered core pipeline with Polars LazyFrames and DuckDB, delivering 3x processing speed and zero-OOM execution on 50M+ row tables.

β€’ Enterprise Scale & Codebase: Scaled platform across 60+ enterprise client projects, expanding core validation engine from 1,000 to 2,560+ LOC.

β€’ Cross-Database Snowflake Pushdown: Staged heterogeneous source data into temporary Snowflake tables, executing distributed joins directly in-warehouse to eliminate multi-terabyte network data transfer.

β€’ Adaptive Memory Allocation: Utilized psutil to evaluate free host RAM at runtime, scaling batch chunks between 70–85% memory capacity with disk-backed chunk streaming (fetch_to_disk).

β€’ Keyless Fallback & Safety: Built automated MD5 hash keyless fallback reconciliation, extraction adapters for MongoDB/DynamoDB document schemas, and automated DML execution blockers to prevent accidental destructive SQL operations.

β€’ GenAI Integration: Decoupled AWS Bedrock / Claude Sonnet diagnostics with pattern clustering, sampling 3 mismatch records per pattern and returning structured JSON root-cause classifications.

DataCraft β€” Autonomous SQL & Data Engineering Assistant

Python β€’ Google Agent Development Kit (ADK) β€’ Gemini API β€’ Firestore β€’ Snowflake β€’ BigQuery β€’ PostgreSQL β€’ Vertex AI

Autonomous multi-modal data engineering agent built with the Google ADK during Google's 4-Hour "Build with Gemini" hackathon, earning the official Google Developer Badge & Credly Certification.

β€’ Dynamic Schema Catalog: Automated Firestore schema registration cataloging database schemas, primary keys, and data types across heterogeneous sources (list_tables, get_table_details, add_table).

β€’ Anti-Pattern Detection: AST parsing with sqlparse to detect full table scans, missing filters, and uncapped sorting with Snowflake, BigQuery, and PostgreSQL optimizations.

β€’ Multi-Modal Architecture Generation: Produced visual ER diagrams via gemini-3.1-flash-lite-image and animated Kafka event-streaming architecture videos via gemini-omni-flash-preview on Vertex AI with GCS storage.

β€’ Vertex AI Memory Bank: Integrated PreloadMemoryTool for session-level dialect persistence, paired with AgentEngineSandboxCodeExecutor and a responsive FastAPI/A2UI card interface.

Enterprise Packaging & Multi-Tier Self-Updater

Python β€’ PyInstaller (onedir/onefile) β€’ Windows API β€’ SQLite β€’ UNC Network Shares

Architected a zero-downtime, atomic hot-swap updater (Launcher → Updater → Real App) distributing updates across 60+ enterprise installations over UNC network shares.

β€’ Atomic Hot-Swap: Decoupled process execution so the updater replaces running binaries without file-lock collisions, backed by rollback protection.

β€’ Zero Console Flashing: Leveraged Windows API process handling to provide silent execution with clean SQLite telemetry logging and UNC share distribution (install_info.json).

β€’ CLI Maintenance Utility: Authored IMIG_Updater.cmd command-line tools for automated environment diagnostics, repairs, and distribution audits.

Healthcare Payer Data Integration Pipeline

Informatica PowerCenter β€’ Oracle SQL β€’ XML / XSD β€’ Flat File Processing

Production ETL integration pipeline processing fixed-width and comma-delimited healthcare feeds into dimensional Oracle tables with automated XML welcome letter distribution.

β€’ Multi-Feed Processing: Cleaned, validated, and normalized multi-tier patient datasets across complex business transformations.

β€’ Schema Compliance: Validated outgoing XML welcome records against strict enterprise XSD schema definitions.

// VERIFIED RECOGNITION

Credentials & Education

Verified industry credentials, certifications, and academic foundations.

πŸ†

Build with Gemini (Track 3) β€” Software Developer

Google Developers β€’ Issued Feb 2026
View Google Badge β†—  β€’  Verify on Credly β†—
πŸŽ–οΈ

Context Engineering Foundation

Cognizant β€’ RAG & AI Systems β€’ Issued Jan 2026
Verify on Credly β†—
πŸŽ“

B.Tech in Computer Science & Engineering (AI)

Noida Institute of Engineering & Technology (NIET), Greater Noida β€’ 2021 – 2025 β€’ CGPA: 7.18
Specialization in Artificial Intelligence, Distributed Systems & Data Engineering
🏫

Intermediate & Matriculation (CBSE)

Jesus & Mary Academy, Darbhanga, Bihar β€’ 2018 – 2021
Science & Mathematics Stream

Ready to Accelerate Your Data Architecture?

Available for Senior Data Engineer, Lead Automation Engineer, and Technical Product Owner opportunities. Let's discuss high-throughput reconciliation, Polars, DuckDB, or GenAI diagnostics.

Copied to clipboard!