Showing posts with label python. Show all posts
Showing posts with label python. Show all posts

Wednesday, May 22, 2024

Calling Python Functions from kdb+/q with PyKX

PyKX allows you to call Python functions from kdb+/q (and vice versa), enabling powerful data analysis using the rich ecosystem of libraries available in Python. In this post, I will show how you can invoke a Lasso Regression function in Python by passing a table from a q script.

1. Install PyKX

pip install pykx

2. Create a Python (.p) file

Create a Python file called lasso.p containing a function that takes a Pandas DataFrame, performs Lasso regression (using the scikit-learn machine learning library), and returns a vector of coefficients.

# lasso.p

import numpy as np
import pandas as pd
from sklearn.linear_model import Lasso

def lasso_regression(df):
    X = df.iloc[:, :-1]
    y = df.iloc[:, -1]

    # Perform Lasso regression
    lasso = Lasso(alpha=0.1)
    lasso.fit(X, y)

    coefficients = np.append(lasso.coef_, lasso.intercept_) 
    return coefficients

3. Invoke the Python function from a q script

Next, write a q script that generates a table of random data and invokes the Python function with it.

// Load pykx and the python file
\l /path/to/python/site-packages/pykx/pykx.q
\l lasso.p

// Create a sample table with random x and y values
n:100;
x:n?10f;
y:2*x+n?2f;
data:([]x;y);

// Call the python function
qfunc:.pykx.get[`lasso_regression;<];
coefficients:qfunc data;

Conversion of data types between kdb+/q and Python

When transferring data between q and Python, PyKX applies "default" type conversions. For instance, tables in q are automatically converted to Pandas DataFrames, and lists are converted to NumPy arrays. You can call .pykx.setdefault to change the default conversion type to Pandas, Numpy, Python, or PyArrow. PyKX also provides functions to convert q data types to specific Python types, such as .pykx.tonp which tags a q object to be converted to a NumPy object. The following code illustrates type conversion:

q) .pykx.util.defaultConv
"default"

// lists are converted to NumPy arrays by default
q) .pykx.print .pykx.eval["lambda x: type(x)"] til 10
<class 'numpy.ndarray'>

// tables are converted to Pandas DataFrames by default
q) .pykx.print .pykx.eval["lambda x: type(x)"] ([] foo:1 2)
<class 'pandas.core.frame.DataFrame'>

// change default conversion to NumPy
q) .pykx.setdefault["Numpy"]

// tables are NumPy arrays now
q) .pykx.print .pykx.eval["lambda x: type(x)"] ([] foo:1 2)
<class 'numpy.recarray'>

// change default conversion to Python
q) .pykx.setdefault["Python"]

// tables are converted to dict when using Python conversion
q) .pykx.print .pykx.eval["lambda x: type(x)"] ([] foo:1 2)
<class 'dict'>

// tag a q object as a Pandas DataFrame
q) .pykx.print .pykx.eval["lambda x: type(x)"] .pykx.topd ([] foo:1 2)
<class 'pandas.core.frame.DataFrame'>

Saturday, December 09, 2023

Python: Running Tasks in Parallel

The concurrent.futures module can be used to run tasks in parallel in Python. Here is an example:

import concurrent.futures
from concurrent.futures import ProcessPoolExecutor

with ProcessPoolExecutor(max_workers = 10) as executor:
    futures = [executor.submit(perform_task, task) for task in tasks]
	
results = [future.result() for future in futures]

The ProcessPoolExecutor class uses a pool of processes to execute tasks asynchronously. The submit function immediately returns a Future object, and you can call future.result(), which will block until the task has completed.

Note that there is also a ThreadPoolExecutor class which uses a pool of threads; however, due to the Global Interpreter Lock, only one thread can execute python bytecode at any one time, which means that you will not achieve any parallelisation in most cases.

Saturday, September 02, 2023

Python: Printing the Stdout/Stderr of a Subprocess

This is how you can run a subprocess in python and print out its stdout and stderr:

import subprocess

proc = subprocess.Popen(["/path/to/myscript", "arg1", "arg2"], 
            stderr=subprocess.STDOUT, stdout=subprocess.PIPE)
for line in proc.stdout:
    print(line.decode().rstrip())
proc.wait()
if proc.returncode != 0:
    print("Command failed with status:", proc.returncode)

Check out the subprocess documentation for more information.

Related post:
Python Cheat Sheet

Thursday, September 02, 2010

Python Cheat Sheet

I've been playing around with python over the last few days. It's really easy to use and I've written the following cheat sheet so that I don't forget how to use it:

Collections:

###################### TUPLE - immutable list
rgb = ("red", "green", "blue")

assert len(rgb) == 3
assert min(rgb) == "blue"
assert max(rgb) == "red"

#find the last element
assert rgb[len(rgb) - 1] == "blue"

assert "red" in rgb
assert "yellow" not in rgb

#slice the list [1,3)
rgb[1:3] #('green', 'blue')

#concatenate twice
rgb * 2

#iterate over list
for e in rgb:
 print e

###################### LIST - mutable list
colors = ["black", "white"]
assert len(colors) == 2

#add an element
colors.append("green")
assert len(colors) == 3
assert colors[2] == "green"

#remove an element
colors.remove("white")
assert "white" not in colors

colors.append("green")

#count occurrences of an element
assert colors.count("green") == 2

#return first index of value
assert colors.index("black") == 0

#insert element
colors.insert(1, "orange")
assert colors.index("orange") == 1

#remove the last element
e = colors.pop()
assert e == "green"

#replace a slice
colors[1:3] = ["red"];

colors.reverse()
colors.sort()

###################### SET
l = ["a", "b", "a", "c"]
assert len(set(l)) == 3

###################### DICTIONARY- mappings

dict = {"alice":22, "bob":20}
assert dict.get("alice") == 22

#add an element
dict["charlie"] = 25
assert dict.get("charlie") == 25

#get list of keys
assert len(dict.keys()) == 3

#get list of values
assert len(dict.values()) == 3

#check for key
assert dict.has_key("alice")
Loops:
for i in range(1, 10, 2):
    print i

while True:
    print 1
    break
Methods:
Printing out the factorial of a number.
def factorial(n):
    """prints factorial of n using recursion."""
    if n == 0:
        return 1
    else:
        return n * factorial(n - 1)

print factorial(4)
Classes:
class Person:

    def __init__(self, name, age):
        self.name = name
        self.age = age

    def equals(self, p):
        return isinstance(p, Person) and \
               self.name == p.name and \
               self.age == p.age


    def toString(self):
        return "name: %s, age: %s" % (self.name, self.age)


p = Person("Bob", 20)
q = Person("Charlie", 20)
assert p.equals(p)
assert p.equals(q)==False
print p.toString()
Read user input:
string = raw_input("Enter your name: ")
number = input("Enter a number:")
File I/O:
#write to file
filename = "C:/temp/test.txt"
f = file(filename, "w")
f.write("HELLO WORLD\nHELLO WORLD")
f.close()

#open a file in read-mode
f = file(filename, "r")

#read all lines into a list in one go
assert len(f.readlines()) == 2

#print out current position
print f.tell()

#go to beginning of file
f.seek(0)

#read file line by line with auto-cleanup
with open("C:\\temp\hello.wsdl") as f:
    for line in f:
        print line
Pickling (Serialisation):
#pickling
f = file("C:\\temp\dictionary.txt", 'w')
dict = {"alice":22, "bob":20}
import pickle
pickle.dump(dict, f)
f.close()

#unpickling
f = file("C:\\temp\dictionary.txt", 'r')
obj = pickle.load(f)
f.close()
assert obj["alice"] == 22
assert obj["bob"] == 20
Exception handling:
try:
    i = int("Hello")
except ValueError:
    i = -1
finally:
    print "finally"


#creating a custom exception class
from repr import repr

class MyException(Exception):
    def __init__(self, value):
        self.value = value
    def __str__(self):
        return repr(self.value)

try:
    raise MyException(2*2)
except MyException as e:
    print 'My exception occurred, value:', e.value
Lambdas:
def make_incrementor (n):
    return lambda x: x + n

f = make_incrementor(2)
print f(42)

print map(lambda x: x % 2, range(10))