News Feed

No results for 'undefined'
Powered by Algolia

Optimize inserting data to Cassandra database through Python driver


Author: GecKo

Originally Sourced from: https://stackoverflow.com/questions/64350512/optimize-inserting-data-to-cassandra-database-through-python-driver

I try to insert 150.000 generated data to the Cassandra using BATCH in Python driver. And it take approximately 30 seconds. What should I do to optimize it and insert data faster ? Here is my code:

from cassandra.cluster import Cluster
from faker import Faker
import time
fake = Faker()

cluster = Cluster(['127.0.0.1'], port=9042)
session = cluster.connect()
session.default_timeout = 150
num = 0
def create_data():
    global num
    BATCH_SIZE = 1500
    BATCH_STMT = 'BEGIN BATCH'

    for i in range(BATCH_SIZE):
        BATCH_STMT +=  f" INSERT INTO tt(id, title) VALUES ('{num}', '{fake.name()}')";
        num += 1

    BATCH_STMT += ' APPLY BATCH;'
    prep_batch = session.prepare(BATCH_STMT)
    return prep_batch

tt = []
session.execute('USE ttest_2')

prep_batch = []
print("Start create data function!")
start = time.time()
for i in range(100):
    prep_batch.append(create_data())

end = time.time()
print("Time for create fake data: ", end - start)

start = time.time()

for i in range(100):
    session.execute(prep_batch[i])
    time.sleep(0.00000001)

end = time.time()
print("Time for execution insert into table: ", end - start)