Cisco logo

Software Engineer - Data & Scalability Platform (8-12 Yrs)

Cisco · Posted Oct 8

Networking, cloud, cybersecurity, and enterprise IT infrastructure services

Bangalore, IndiaFull-timeHybridLead/Staff8–12 years₹45.0L–₹70.0L yearly100+ applicants
NetworkingCybersecurityCloud ComputingInformation TechnologyHardwarePublic Company
Full time

Get a personal compatibility score

Add a resume for personal matches

About the role

The Splunk Agent Resilience team at Cisco provides scalable, cost-effective evaluation and guardrails that ensure AI agents behave as intended, improving reliability and reducing risks for customers deploying AI-powered applications. As a Software Engineer on the Data & Scalability Platform team, the role owns the data plane end to end — relational, analytical/columnar, object storage, caches, and queues — along with the streaming and compute components that move and transform data and the performance tooling that validates behavior under load. The engineer will make this data plane correct, fast, cost-bounded, and tenant-isolated at production scale, own the capacity and performance model, and partner closely with the product engineering teams building on top of it.

What you will do

  • Own and evolve the platform's data stores — relational, analytical/columnar, object storage, and caching — including schema design, migration safety, and retention.
  • Design and operate the streaming and queueing backbone that carries data from ingest to queryable.
  • Build and scale the compute and pipelines that move and transform data, including stream processors, writers, schedulers, and distributed worker fleets.
  • Own end-to-end data performance: query and write latency, indexing and sharding strategy, and hot-path optimization.
  • Own the capacity model for the data platform, along with quotas, rate limiting, and multi-tenant isolation.
  • Drive cost efficiency across the data platform, including cost per unit of telemetry and cost attribution.
  • Build and own performance tooling — load testing, profiling, and benchmarking — used to validate scale and guide optimization.
  • Build observability instrumentation and telemetry pipelines so platform signals are correct, complete, and affordable.
  • Own data operations: backup and restore, retention and deletion, replication, migration and backfill, and data-quality signals.
  • Design, develop, test, and maintain production services and internal tooling in Python and/or Go, using secure coding practices and automated tests.
  • Debug and resolve complex production issues across data stores, queues, and services, and contribute to monitoring, on-call support, and root-cause analysis.
  • Contribute to platform architecture direction and act as a technical resource and mentor through design reviews, code reviews, and documentation.

Skills used in this role

PythonGoJavaC++PostgreSQLMySQLClickHouseDruidBigQuerySnowflakeKafkaRabbitMQPulsarKinesisKubernetesAWSGCPOpenTelemetryPyroscopepy-spyCeleryTemporalFlinkSpark StreamingTrino/Presto

What the employer is looking for

  • Bachelor's degree with 7+ years of related experience, or Master's degree with 4+ years, or PhD with 1+ year of related experience, in Computer Science, Software Engineering, or a related field.
  • Strong backend software engineering experience building, operating, and delivering highly scalable, reliable, production-grade data-intensive services and platforms.
  • Proven experience with large-scale distributed systems, including designing and solving complex problems of scalability, availability, performance, reliability, and fault tolerance.
  • Hands-on experience operating and tuning both an OLTP database (e.g., PostgreSQL, MySQL) and an analytical, columnar, or time-series store (e.g., ClickHouse, Druid, BigQuery, Snowflake) — schema design, indexing, query tuning, and migration safety.
  • Production experience with streaming or queueing systems (e.g., Kafka, RabbitMQ, Pulsar, Kinesis) — partitioning, consumer-group semantics, delivery guarantees, and backlog and dead-letter handling.
  • Experience scaling data-processing compute — queue consumers, async task workers, or streaming jobs — including concurrency tuning, batching, backpressure, and throughput behavior under sustained load.
  • Strong proficiency in Python and solid experience with at least one additional backend programming language such as Go, Java, C++, or similar.
  • Demonstrated performance engineering ability — profiling, benchmarking, and load testing real systems, and turning the measurements into capacity, design, and cost decisions.
  • Experience with cloud-native technologies, containerized environments (Kubernetes), and public cloud platforms (AWS, GCP, or similar).
  • Demonstrated ability to independently design, develop, debug, test, and maintain software with minimal guidance.

Preferred qualifications

  • Experience running ClickHouse or a comparable columnar store at scale — sharding and replication topology, materialized views, merge and mutation behavior, retention and TTL mechanics.
  • Experience owning a capacity model or a FinOps practice: cost per unit of telemetry, cardinality governance, and quota or rate-limit design.
  • Experience with multi-tenant isolation and noisy-neighbor mitigation in shared data systems.
  • Experience with backup and restore, tested disaster-recovery drills against stated RPO/RTO, and data retention or deletion compliance (e.g., GDPR/DSR hard delete).
  • Experience building observability instrumentation and pipelines (OpenTelemetry, collectors, metrics/tracing backends) rather than only consuming dashboards.
  • Experience building or operating performance tooling as a product for other engineers — load-test harnesses, continuous profiling (e.g., Pyroscope, py-spy, flame-graph workflows), or benchmarking suites used to justify capacity and design decisions.
  • Experience with distributed task frameworks and worker orchestration (e.g., Celery, Temporal, Flink, custom consumer fleets) at production scale.
  • Familiarity with streaming databases, stream processing frameworks, or distributed query engines (e.g., RisingWave, Flink, Spark Streaming, Trino/Presto) and with model-serving infrastructure.
  • Experience supporting stateful systems in both cloud/SaaS and air-gapped or on-prem deployments.
  • Experience leading medium-sized features or projects end-to-end, mentoring engineers, and influencing technical decisions across teams.
  • Strong communication skills, with the ability to turn complex performance, capacity, and cost findings into clear recommendations for engineering and product partners.

Benefits and support

  • Compensation and benefits are detailed in the job posting

About Cisco

Cisco Systems, Inc. is a worldwide leader in IT, networking, and cybersecurity solutions that transform how people connect, communicate, and collaborate. The company specializes in routing, switching, data center infrastructure, and software-defined networking designed for enterprises of all sizes. Today, Cisco drives major advancements in AI infrastructure, observability, and secure hybrid work platforms.

Industry
Networking
Company size
90000+ employees
Founded
1984
Location
San Jose, California, USA
Funding stage
Public Company

Leadership

CR
Chuck Robbins

Chair and Chief Executive Officer

JP
Jeetu Patel

President and Chief Product Officer

MP
Mark Patterson

Executive Vice President and Chief Financial Officer

FK
Francine Katsoudas

Executive Vice President and Chief People, Policy & Purpose Officer

LC
Liz Centoni

Executive Vice President and Chief Customer Experience Officer